AT A GLANCE
Backend/Data Engineer II building AI-augmented, self-healing data pipelines and streaming systems at scale on GCP
KEY SKILLS
PythonScalaJavaSQLApache SparkApache BeamApache FlinkKafkaBigQueryDataflowApache AirflowKubernetesTerraformdbt
WHAT YOU'LL DO
- Design and build self-healing, schema-validated data pipelines on GCP (BigQuery, Dataflow, Pub/Sub)
- Implement and optimize large-scale processing workflows using Apache Spark, Flink, or Beam
- Own real-time Kafka consumer pipelines and scheduled batch jobs with sub-hour SLAs
- Implement Schema Registry and data-contract standards across producers and consumers
- Build and maintain feature stores and data pipelines for production ML models
- Champion data cataloguing, lineage tracking, and access-control policies
- Monitor and optimize BigQuery slot usage and pipeline costs
- Collaborate cross-functionally with Data Science, DevOps, and Security teams