About the Role
Design, develop, deploy and maintain reliable ETL/ELT batch & streaming data pipelines, ingest structured/semi-structured/unstructured data from multiple source systems (DB, API, log, Kafka etc.) Perform dimensional data modeling, build and iterate data warehouse / data lake, define table schema, partition strategy, data layers (ODS/DWD/DWS/ADS)51job Write & optimize complex SQL, tune query performance, reduce resource cost and improve pipeline stability and SLA Implement data quality rules, anomaly detection, monitoring & alerting, troubleshoot data lineage and data consistency issues Use workflow orchestration tools to schedule, manage and monitor data jobs (Airflow/Dagster) Collaborate with analysts, DS and business stakeholders to translate business requirements into data solutions, build reusable datasets and data APIs Participate in data governance: metadata management, data security, access control, documentation for schema and pipelines51job Evaluate and adopt cloud data technologies (Snowflake/BigQuery/Redshift), continuously optimize storage and compute cost
Experience on AWS/Azure/GCP cloud data stack Real-time streaming data pipeline development (Kafka, Flink, Pulsar) Data governance, data lineage, metadata platform construction experience Experience supporting ML feature platform / feature engineering English working proficiency Domain knowledge: automotive, manufacturing, finance, retail etc.
Internal Referral bonus of this vacancy: RMB 3000(Valid only for Bosch associates).For the detailed regulation. please refer to Bosch China Internal Referral Policy