Skip to content
Computese home

08

Data platform

Ingestion, a governed warehouse or lakehouse and a semantic layer, engineered with tests, lineage and access control, so every team reads the same number and your own staff can run the platform after we leave.

  1. Sources
  2. Ingest
  3. Model
  4. Serve
PostgreSQLApache KafkadbtPower BI
Start with
Data platform assessment
Ways to engage
Dedicated team, time and materials, fixed price
Works to
DAMA-DMBOK, Data contracts, Medallion architecture
Reply time
Within 24 hours

Who it is for.

If you run a business

Your reports come from spreadsheets stitched together by hand, and nobody trusts the totals.

Sound familiar?

  • Two departments bring different revenue numbers to the same meeting.
  • A monthly report takes days of copying and pasting.
  • Decisions wait for someone to pull the data.
  • Nobody is sure who can see customer data.

If you lead a technology team

You need ingestion, modelling and serving built like software: version control, CI, data contracts, lineage and fine-grained access, on a platform your team can operate.

Sound familiar?

  • Pipelines break silently and are noticed by the people reading the dashboard.
  • Transformations live in untested SQL scripts and BI tools.
  • There is no lineage, so an upstream change breaks something unknown.
  • Warehouse cost grows faster than usage.

What changes.

What you can hold the work to, in plain terms.

  1. 01

    One definition of every number.

    Metrics defined once in a semantic layer and used by every dashboard, API and notebook.

  2. 02

    Pipelines that fail loudly.

    Tests, freshness checks and data contracts stop bad data before it reaches a report.

  3. 03

    Governed by design.

    A catalog, lineage and role-based access, with personal data classified and protected.

  4. 04

    A platform your team owns.

    Code in Git, CI for every change, documentation and training, so the platform outlives the project.

A lakehouse with governance built in.

Data moves through layers of increasing trust, and governance runs underneath every one of them.

Sources (databases through change data capture, SaaS APIs, files and events) are ingested in batch or streaming with schema checks. The lakehouse holds bronze (raw, append-only, replayable), silver (cleaned, typed, conformed) and gold (business-ready models) layers, governed by a catalog with lineage, tests and contracts, access policies, and freshness and cost monitoring. Data is served to dashboards, data APIs and machine-learning features. Orchestration uses Airflow or Dagster schedules, dbt builds with tests, CI for every model change and a semantic layer for shared metrics.

Fig. 1 Reference architecture for a data platform. Warehouse or lakehouse, cloud or on-premises: the layers and the controls stay the same.
Close-up of hot-swap NVMe drive sleds in a storage server, one activity light glowing orange.

What we bring.

The disciplines inside this service, and the detail we work to in each.

  • 01

    Data strategy and modelling

    The questions the business needs answered, turned into a model that can answer them, and a roadmap ordered by value.

    • Source inventory and ownership
    • Business glossary and metric definitions
    • Dimensional and data vault modelling
    • Roadmap by value and effort
  • 02

    Ingestion and streaming

    Reliable movement from every source, in batch or in real time.

    • Change data capture with Debezium
    • Apache Kafka and managed streams
    • Managed connectors where they fit
    • Schema evolution handled, not ignored
  • 03

    Warehouse and lakehouse

    Storage and compute chosen for your volumes, skills and budget.

    • Snowflake, BigQuery, Databricks or PostgreSQL
    • Apache Iceberg and Delta Lake tables
    • Medallion layers: bronze, silver, gold
    • Partitioning and cost controls
  • 04

    Analytics engineering

    Transformations written as tested, reviewed, versioned code.

    • dbt models with tests
    • Data contracts with producing teams
    • CI for every model change
    • Documentation generated from code
  • 05

    Governance and security

    Know what data exists, where it came from and who may see it.

    • Catalog and column-level lineage
    • Role, row and column-level access
    • PII classification and masking
    • Retention rules and audit trails
  • 06

    Serving: BI, APIs and AI

    Trusted data delivered where decisions are made, from the boardroom dashboard to the feature store behind a model.

    • A semantic layer for shared metrics
    • Power BI, Looker or Metabase dashboards
    • Data APIs for applications
    • Features and datasets for ML and AI
    • Freshness and quality targets per dataset
    • Self-service within guardrails

Which platform, and why.

The platform is chosen after we have seen your sources, your volumes and your team, not before. If what you run today is sound, we build on it.

PlatformFits whenStrengthWatch for
PostgreSQLFits whenVolumes are modest and the team already knows SQL.StrengthLow cost, simple to run, one familiar engine.Watch forHeavy analytical queries need their own replica or instance.
SnowflakeFits whenA SQL-first team wants elastic compute with little administration.StrengthStorage and compute scale separately, per workload.Watch forCredit spend needs budgets, auto-suspend and monitoring.
DatabricksFits whenLarge volumes, streaming and data science share the same data.StrengthDelta Lake tables, Spark and notebooks in one platform.Watch forMore engineering depth is needed to run it well.
BigQueryFits whenYou run on Google Cloud, or workloads are bursty.StrengthServerless: no clusters to size or patch.Watch forOn-demand cost follows bytes scanned, so partitioning and clustering matter.

Whichever platform it is, models live in version control beside their tests, so the business logic is never locked inside one vendor.

How it runs.

Every stage ends with a document you keep and a gate you can check.

  1. 01

    Map

    Sources, owners and the definitions that matter.

    Exit gate: Every key metric has one definition and one owner.

    • Source inventory
    • Metric glossary
    • Access model

    You receiveData model and definitions

  2. 02

    Build

    Pipelines and models, tested and in version control.

    Exit gate: Models pass their tests in CI before they reach production.

    • Ingestion
    • dbt models
    • Data tests

    You receivePipelines in production

  3. 03

    Serve

    Dashboards, APIs and a semantic layer for each team.

    Exit gate: New numbers reconciled with the old ones and signed off by the business.

    • Semantic layer
    • Dashboards
    • Data APIs

    You receiveDashboards and data APIs

  4. 04

    Hand over

    Your team runs it, with our documentation and training.

    Exit gate: Your team has shipped a change on its own, with us alongside.

    • Runbooks
    • Training
    • Support window

    You receiveDocumentation and training

What is in scope.

Written down before work starts, so nothing is assumed.

Included

  • Source inventory, data model and metric definitions
  • Ingestion pipelines, batch or streaming
  • Warehouse or lakehouse set-up
  • Tested transformations with documented definitions
  • Dashboards, a semantic layer and data APIs
  • Access control, lineage, documentation and handover

Not included

  • Manual data entry or clean-up inside source systems
  • Platform licences, billed to you by the vendor

Standards and stack.

The public frameworks we measure the work against, and the platforms we run it on.

Standards we work to

DAMA-DMBOK
Data management practice: ownership, quality, metadata and governance.
Data contracts
Schemas and expectations agreed with the teams that produce data.
Medallion architecture
Bronze, silver and gold layers of increasing trust.
OpenLineage
Lineage captured in an open standard, not locked into one tool.
Privacy by design
Personal data classified, minimised and access-controlled to the laws that apply to you.

How we choose tools

Certified engineers
AWS Solutions Architect, Azure Solutions Architect Expert, Google Cloud and security certifications, held by the engineers who do the work.
Licensed tools only
Every tool comes from an approved list: commercial software under its licence, or open source under a standard licence. Nothing cracked, nothing unlicensed.
Your platform first
Where you already run something that works, we build on it.
Not on the list?
Ask. Engineers who know the fundamentals pick up a new tool quickly, and we will tell you plainly if we have not used it before.

Platforms and tools we work with

Warehouses and lakehouses

  • PostgreSQL
  • Snowflake
  • Databricks
  • Google BigQuery
  • Amazon Redshift
  • Azure Synapse Analytics
  • Microsoft Fabric
  • ClickHouse
  • DuckDB

Table formats and storage

  • Apache Iceberg
  • Delta Lake
  • Apache Parquet
  • Amazon S3
  • Azure Data Lake Storage
  • Google Cloud Storage
  • MinIO

Ingestion and change data capture

  • Apache Kafka
  • Confluent
  • Kafka Connect
  • Debezium
  • Airbyte
  • Fivetran
  • dlt
  • AWS DMS
  • Azure Data Factory
  • AWS Glue
  • Apache NiFi

Transformation

  • dbt
  • SQLMesh
  • Apache Spark
  • Apache Flink
  • Polars
  • pandas
  • SQL
  • Python

Orchestration

  • Apache Airflow
  • Dagster
  • Prefect
  • AWS Step Functions

Quality, catalogue and lineage

  • dbt tests
  • Great Expectations
  • Soda
  • OpenLineage
  • DataHub
  • OpenMetadata
  • Unity Catalog
  • Microsoft Purview
  • AWS Lake Formation

Semantic layer and BI

  • dbt Semantic Layer
  • Cube
  • Power BI
  • Tableau
  • Looker
  • Looker Studio
  • Metabase
  • Apache Superset
  • Grafana

Machine learning and AI

  • MLflow
  • Amazon SageMaker
  • Vertex AI
  • Azure Machine Learning
  • pgvector
  • Feast

Where we have done it.

Client cases name the industry and the stack, never the client.

How to start.

A fixed, small first engagement, then the model that fits the rest.

A first engagement

Data platform assessment

A review of your sources, reports and current pipelines, ending in a target architecture and a first use case worth building.

You receive

  • Inventory of sources and reports
  • Where the numbers disagree, and why
  • Target architecture and tool choice
  • A first use case, scoped and priced

What we need from you

  • Read access to the main sources
  • A named owner for each business area
  • The two or three reports that matter most
  • The numbers that currently disagree

Then, the model that fits

  • Dedicated team

    Long programs such as a platform migration

    A monthly rate per engineer

  • Time and materials

    Ongoing improvement, support and discovery work

    Billed for the time used

  • Fixed price

    Defined projects: a website, an assessment, a migration stage

    One price for that scope

Common questions.

Warehouse or lakehouse?

A warehouse suits structured reporting at modest volume; a lakehouse suits mixed data, larger volumes and machine learning on the same storage. Many teams start with a warehouse and grow into open table formats. The assessment recommends one.

Can you work with the tools we already have?

Yes. Existing warehouses, BI tools and pipelines are kept where they work; tests, lineage and structure are added around them before anything is replaced.

Is the platform ready for AI?

AI needs what analytics needs: clean, documented, governed data. The gold layer and the semantic layer serve dashboards and models alike.

Who owns the platform after the project?

Your team. The code sits in your repositories, and the handover includes runbooks and training.

How do you prove the new numbers are right?

Old and new run side by side. Row counts, totals of key measures and each key metric are reconciled table by table, within a tolerance agreed with you, and the business owner signs off each area before the old report is retired.

Do we have to migrate at all?

Often not. If your current warehouse is sound, we add tests, definitions and monitoring to it rather than moving you.

Guides from the blog.

Plain-language articles on data platform, with their sources.

Data16 min read

Data compression algorithms explained: lossless, lossy and how to choose

Data compression algorithms remove redundancy to store data in fewer bits. How gzip, Brotli, Zstandard, LZ4 and lossy codecs work, and how to choose.

Updated

Data19 min read

Generative AI for databases: text-to-SQL, vector search and safe use

Generative AI for databases turns plain-language questions into SQL and searches data by meaning. How it works, how accurate it is and how to use it safely.

Updated

Data15 min read

Big data analytics explained: how it works, tools, examples and costs

Big data analytics is analyzing data too large, fast or varied for one ordinary database. How it works, the main tools, the four types, costs and governance.

Updated

Data14 min read

The role of IT in healthcare: systems, interoperability and security

How IT runs healthcare: EHR, PACS, portals and telehealth, HL7 FHIR interoperability, the US and EU data-sharing rules, and the security that matters.

Updated

Data18 min read

Building data analytics software: architecture, build vs buy and delivery

How to build data analytics software: ingestion, a warehouse or lakehouse, tested models, a semantic layer, secure embedded dashboards and APIs, in stages.

Updated

All Data articles →

Start with a conversation.

Tell us what you run and what is getting in the way. You get a reply within 24 hours.

Hours
Mon–Fri, 9:00–17:00 ET
Closed on statutory holidays
Office
110 Place d'Orléans Dr
Ottawa, ON K1C 2L9