Skip to main content

Synthetic Data Engineer (AI Data/Training)

Expired
This role has expired and is no longer accepting applications. Browse similar roles →
Hyphen Connect Limited
Australia
remote
Full Time / Permanent

Apply for this job

Posted 26d ago
This role is expired

These roles are hiring now

View all similar roles →

Senior Data Engineer

Octopus Deploy
Australia
remote
$145,000 - $165,000 per yr
  • Build and manage data platform capabilities for internal org and customers
  • 3+ years data or software engineering experience
  • SQL, Python, cloud platforms (Azure/AWS), Snowflake, Databricks, Kafka
Posted 29d ago

Senior Data Engineer - Modelling

Xero
Melbourne, VIC
hybrid
  • Build data pipelines, models and infrastructure at scale
  • Senior-level experience with modern data platforms
  • Snowflake, dbt, Prefect, Terraform, AWS
Posted 8d ago

Staff Data Engineer

Checkbox Technology
Sydney, NSW
hybrid
  • Own data strategy & architecture for AI-native legal SaaS platform
  • Significant senior data engineering or staff engineer experience required
  • AI data patterns, multi-tenant SaaS, retrieval systems, context engineering
Posted 10d ago

Data Engineer

IAG
Sydney, NSW
hybrid
  • Build scalable data pipelines, products, APIs using BigQuery and DBT
  • 4+ years Data Engineering or Software Engineering experience
  • SQL, BigQuery, DBT, GCP, AI automation, DevOps practices
Posted 10d ago

We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.

Responsibilities:

  • Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.
  • Implement automated quality scoring and de-duplication systems.
  • Manage data pipelines that feed directly into SFT and DPO training loops.

Qualifications:

  • Proven experience building large-scale data pipelines (Airflow, Spark, Ray).
  • Deep knowledge of prompt engineering for data generation.
  • Familiarity with dataset distillation and bias mitigation.