Описание: Statista is a global business data platform that provides reliable, easy-to-use data, analytics products, and services to support fact-based decision-making worldwide. Its Healthcare Platform transforms international hospital data into structured, queryable healthcare data assets.
Задачи:
Build, optimize, and operate reliable ELT pipelines using Python, SQL, and Prefect/Airflow; Ingest data from heterogeneous international sources, APIs, databases, and lakehouse storage including S3 and Apache Iceberg; Drive entity resolution and master data management for international hospital entities; Map raw source data to canonical structures and maintain standardized vocabularies such as ICD/OPS and specialty taxonomies; Establish data contracts and schema management with Pydantic and dbt contracts; Ensure dataset reproducibility, data lineage tracking, and automated validation across platform pipelines; Optimize data storage, query execution, and compute costs across AWS and Snowflake; Implement automated testing and deployment workflows for data pipelines using GitHub Actions and Terraform; Partner with Analytics Engineers, Data Scientists, and domain experts to deliver documented, research-grade, production-ready datasets.
Требования:
Advanced Python and analytical SQL for complex data ingestion across diverse file formats, REST APIs, databases, and cloud lakes; Hands-on experience with Prefect, Airflow, or Dagster; Practical experience with entity resolution or record linkage frameworks; Experience with schema management and data contracts using Pydantic, dbt contracts, or JSON Schema; Deep hands-on experience in an AWS production environment, including S3 and ECS/EC2; Experience with cloud data warehouses, including Snowflake; Proven experience automating data pipeline deployments, integration tests, and validation workflows via GitHub Actions; 3+ Years of experience in data engineering building production pipelines and data platforms; Experience taking a core platform or product through build, launch, and iteration within one organization; Bachelor's or Master's degree in Computer Science, Data Science, Software Engineering, or a related quantitative field; Strong analytical and systems mindset; Ability to transform messy, heterogeneous international data into clean, well-governed, highly structured data assets; Fluent English; Highly structured, curious, detail-oriented, and collaborative working style; Nice to have: Knowledge of ontology or semantic frameworks, international medical classifications or vocabularies, metadata registries and lineage catalogs, Terraform, healthcare domain experience, German.
Условия:
Work from abroad up to 30 calendar days a year; Hybrid work and flex-time; International team and social events; Subsidized urban mobility and access to fitness and wellness options; Free access to Langdock; Career and training opportunities; Attractive locations and modern offices; Mental health support with OpenUp; Some benefits apply only to the German entity and to Junior-level roles or above.
#J-18808-Ljbffr
data engineer for healthcare data in Berlin Arbeitgeber: Statista Inc.
Statista ist ein hervorragender Arbeitgeber, der seinen Mitarbeitern nicht nur ein dynamisches und internationales Arbeitsumfeld bietet, sondern auch zahlreiche Möglichkeiten zur beruflichen Weiterentwicklung. Mit flexiblen Arbeitszeiten, der Option auf hybrides Arbeiten und einem starken Fokus auf Teamkultur und Diversität, fühlen sich unsere Mitarbeiter wertgeschätzt und unterstützt. Darüber hinaus profitieren Sie von attraktiven Standorten, modernen Büros und umfassenden Gesundheitsangeboten, die das Wohlbefinden unserer Mitarbeiter fördern.