Data Architect

Data Architect

Vollzeit Homeoffice möglich
H

Senior Data Architect

FullTIme

Remote (US)

Required Experience : Data Platform, Integration, Azure Databricks & Lakehouse Specialist

Professional Summary

Strong hands-on background in PySpark , Medallion Architecture , Delta Live Tables , Unity Catalog , and MongoDB Atlas integration. Familiar with Databricks Genie for AI-assisted analytics. An effective technical leader who sets architectural direction, mentors teams, and drives delivery. Exposure to Power BI and reporting solutions is a plus, not a core requirement.

Core Skills

  • Databricks: Medallion Architecture (Bronze/Silver/Gold), Delta Lake, Delta Live Tables (DLT), Auto Loader, Unity Catalog, Databricks Workflows, Databricks Apps, Databricks Genie
  • Languages: PySpark (expert), Python (PEP 8/PEP 20), SQL
  • Data Platforms: Azure Databricks, ADLS, Azure Synapse, Snowflake
  • Databases & Integration: MongoDB Atlas, SQL Server (JDBC), REST APIs, OAuth (SharePoint), cloudfiles
  • Governance: Unity Catalog - metastores, catalogs, schemas, RBAC, lineage, access policies
  • DevOps: Git, CI/CD for Databricks workflows and DLT pipelines, JSON/YAML config management
  • Reporting (nice-to-have): Power BI, DAX, Power Platform

Medallion Architecture & Databricks Pipelines

  • Designed and implemented end-to-end Lakehouse solutions on Azure Databricks across Bronze/Silver/Gold layers with schema enforcement and data quality checks
  • Delivered production-grade DLT pipelines and Auto Loader streaming ingestion from ADLS and external sources
  • Optimised PySpark jobs for performance and cost - partition tuning, caching, modular function-based code design

Unity Catalog & Governance

  • Implemented Unity Catalog as the foundational governance layer - metastores, catalogs, schemas, fine-grained RBAC, lineage, and audit controls
  • Standardised metadata management and data discovery across the platform

Data Architecture & Integration

  • Architected data flows from MongoDB Atlas, JDBC (SQL Server), REST APIs, and OAuth sources into the enterprise data platform
  • Established architectural standards for ingestion, transformation, and consumption layers; led architectural reviews across squads
  • Built configuration-driven (JSON/YAML) pipeline frameworks enabling scalable onboarding of new data sources

Databricks Genie

  • Familiar with Databricks Genie for enabling natural language querying of data assets and AI-assisted analytics for business users

Leadership

  • Mentored engineers on Databricks, PySpark, and Python best practices (PEP 8/PEP 20, naming conventions, modular design)
  • Conducted code reviews, enforced coding standards, and driven CI/CD adoption for Databricks workflows and DLT pipelines

#J-18808-Ljbffr

H

Kontaktdaten:

Hexaware Technologies, Inc Recruiting-Team