Job summary.
We are building an on-premises analytical data platform to bring together clinical and
operational data from multiple hospital sites into a governed lakehouse (raw, curated, and
serving layers). This role focuses on pipelines, data reliability, and platform operations —
turning captured data into trusted, linked, analysis-ready datasets.
You will work as part of a data team under established architecture and quality standards. A
platform steering committee and senior consultant provide design direction, phase planning,
and review; you will implement, document, and operate the ingestion and transformation layer
with growing independence over time.
Key responsibilities
Data ingestion & platform pipelines
Design, build, and maintain batch and change-capture ingestion from operational databases
across multiple sites.
Implement orchestrated workflows for scheduled extraction, landing, transformation, and
promotion between lakehouse layers.
Ensure pipelines are idempotent, observable, and recoverable (retries, checkpoints, replay
where appropriate).
Work with database administrators on read-only access, change-log readiness, and safe
extract windows.
Lakehouse layers (raw → curated)
Manage immutable raw landing and curated (silver) datasets following medallion-style
layering.
Implement validation, cleansing, typing, and conformed models using transformation-as
code practices.
Support slowly changing history for key demographic and reference entities where attributes
change over time.
Apply configuration-driven pipeline definitions (declarative specs reviewed in version
control) rather than one-off scripts per table.
Cross-site identity & data linking
Implement logic to unify records across sites (e.g. patients, providers, facilities) using
defined matching rules and steward review for ambiguous cases.
Maintain bridge and reference structures that map source identifiers to enterprise identifiers
with full audit trail.
Data quality, reconciliation & operations
Run and automate reconciliation between source systems and platform copies (counts, keys,
samples, freshness).
Respond to pipeline failures and data drift using runbooks and escalation paths.
Contribute to schema change handling (additive vs breaking changes) with documentation
and alerts.
Monitor service levels (lag, success rate, data freshness) via operational dashboards.
Documentation & collaboration
Maintain technical runbooks, pipeline documentation, and change records on the same day
as changes.
Partner with the Analytics Engineer on catalog entries, data dictionary fields, and lineage
metadata.
Support clinical and operational data stewards with technical fixes; business decisions on
duplicates and definitions stay with stewards.