Illustrative engagement — describes the type of work we deliver; client details are withheld.
Business challenge
Legacy data systems and slow reporting. Finance and risk teams waited until mid-morning for the previous day's numbers, and month-end reconciliations depended on manual spreadsheet work.
Existing environment
- On-premise relational warehouse near capacity
- Hundreds of stored procedures run as nightly batches
- Reports built separately by each department
- No lineage or data quality monitoring
Technical challenges
- Untangling undocumented transformation logic
- Migrating without disrupting regulatory reports
- Establishing access controls for sensitive financial data
Our approach
- 1Catalogued every source, job and report consumer
- 2Designed a bronze–silver–gold lakehouse on Azure
- 3Rebuilt transformations in dbt with automated tests
- 4Ran old and new pipelines in parallel and reconciled outputs before cut-over
Architecture overview
- IngestionChange-data-capture from core systems into cloud storage
- ProcessingDatabricks / Spark jobs orchestrated by workflow scheduler
- Modellingdbt models with tests, documentation and lineage
- ConsumptionCertified Power BI datasets and self-service reports
Technologies used
- Databricks
- Azure
- Spark
- dbt
- Power BI
Implementation
- Phase 1
Discover
Source and report inventory, prioritised migration backlog
- Phase 2
Foundation
Landing zone, security model, CI/CD for data code
- Phase 3
Migrate
Subject-area waves with parallel runs and reconciliation
- Phase 4
Optimise
Cost tuning, decommissioning of legacy jobs, team enablement
Results
- Improved data availability for morning reporting
- Reduced batch processing time
- Automated data quality checks on critical tables
- A single certified dataset per business domain
Business impact
Finance and risk teams begin the day with current numbers, and the platform is ready to support forecasting and machine learning use cases.