- Design, develop, and maintain scalable ETL/ELT data pipelines
- Build cloud-based data solutions using GCP and AWS services
- Use Databricks, Apache Spark, and PySpark for large-scale data processing
- Integrate structured, semi-structured, and unstructured data from multiple sources
- Develop and optimize batch and real-time data-processing workflows
- Improve pipeline performance, reliability, scalability, and cost efficiency
- Implement data-quality checks, monitoring, security, and governance standards
- Design and support cloud data warehouses and data lakes
- Troubleshoot production issues and perform root-cause analysis
- Collaborate with data architects, analysts, application teams, and business stakeholders
- Create and maintain technical documentation for pipelines, data models, and workflows
Requirements
- 6+ Years experience required in the position overview
- 5+ years of professional data engineering experience
- Strong hands-on experience with both GCP and AWS
- Expertise in Databricks, Apache Spark, and PySpark
- Strong programming skills in Python and SQL
- Proven experience developing ETL/ELT pipelines and cloud data platforms
- Experience with data warehouses, data lakes, and dimensional data modeling
- Experience with orchestration tools such as Apache Airflow or Google Cloud Composer
- Understanding of data security, governance, monitoring, and quality frameworks
- Strong analytical, troubleshooting, and problem-solving skills
- Excellent communication and cross-functional collaboration skills
- Preferred: experience with BigQuery, Cloud Storage, Dataflow, Dataproc, Pub/Sub, Cloud Composer, S3, Glue, EMR, Redshift, Lambda, Kinesis, Delta Lake, Databricks Lakehouse, Apache Kafka, CI/CD, Git, Terraform, cloud infrastructure automation, and Agile development environments
Core Competencies
Demonstrates expertise in designing and developing scalable ETL/ELT data pipelines using GCP and AWS, with strong proficiency in Databricks, Apache Spark, and PySpark. Capable of optimizing data workflows and ensuring data quality, security, and governance across cloud data solutions.
Highest-signal resume keywords
- ETL/ELT Pipeline Development
- GCP and AWS Expertise
- Databricks, Apache Spark, and PySpark
- Data Warehouse and Data Lake Design
- Data Security and Governance
Hard Skills
- Python
- SQL
- Data Engineering
- Data Modeling
- Batch and Real-Time Processing
- Data Quality Checks
- Root-Cause Analysis
- Orchestration Tools
- Cloud Data Solutions
- Performance Optimization
Soft Skills
- Analytical Skills
- Troubleshooting Skills
- Problem-Solving Skills
- Communication Skills
- Cross-Functional Collaboration
Industry Keywords
- Cloud Infrastructure Automation
- Agile Development
- Data Governance
- Data Security
- Data Monitoring
Tools & Technologies
- Apache Airflow
- Google Cloud Composer
- BigQuery
- Cloud Storage
- Dataflow
- Dataproc
- Pub/Sub
- S3
- Glue
- EMR
#J-18808-Ljbffr
Data Engineer – GCP, AWS, Databricks Arbeitgeber: Jobtailor
Als Front Office Supervisor in unserem dynamischen Team bieten wir Ihnen die Möglichkeit, in einem unterstützenden und freundlichen Arbeitsumfeld zu wachsen. Wir legen großen Wert auf die berufliche Entwicklung unserer Mitarbeiter und bieten regelmäßige Schulungen sowie die Chance, Verantwortung zu übernehmen. Unsere Lage ermöglicht es Ihnen, Teil einer lebendigen Gemeinschaft zu sein, während Sie gleichzeitig die Standards unseres Franchise-Partners einhalten und unseren Gästen einen unvergesslichen Aufenthalt bieten.