Skip to main content

Synthetic Data Engineer (AI Data/Training)

Expired
This role has expired and is no longer accepting applications. Browse similar roles →
Hyphen Connect Limited
Australia
remote
Full Time / Permanent

Apply for this job

Posted 1 month ago
This role is expired

These roles are hiring now

View all similar roles →

Senior Data Engineer (Python / AWS / ML Pipelines)

Jobgether
Australia
remote
  • Build and operate large-scale data and ML pipelines in production
  • Senior-level experience with production data and ML systems
  • Python, AWS (SageMaker, Glue, Step Functions), Apache Airflow
Posted 4d ago

Senior Data Engineer

Octopus Deploy
Australia
remote
$145,000 - $165,000 per yr
  • Build and manage data platform capabilities for internal org and customers
  • 3+ years data or software engineering experience
  • SQL, Python, cloud platforms (Azure/AWS), Snowflake, Databricks, Kafka
Posted 1 month ago

Lead Data Engineer

Monash Health
Clayton, VIC
  • Build modern data pipelines and AI platforms for healthcare
  • Significant experience in modern lakehouse architectures and data engineering
  • Python, SQL, Azure, Fabric, CI/CD, LLMs, data governance
Posted 6d ago

AI Data Engineer

Wesfarmers
Melbourne, VIC
hybrid
  • Design and build data components for AI and GenAI solutions
  • 5+ years data engineering experience
  • Snowflake, Microsoft Fabric, RAG, embeddings, vector search, SQL, Python
Posted 14d ago

We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.

Responsibilities:

  • Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.
  • Implement automated quality scoring and de-duplication systems.
  • Manage data pipelines that feed directly into SFT and DPO training loops.

Qualifications:

  • Proven experience building large-scale data pipelines (Airflow, Spark, Ray).
  • Deep knowledge of prompt engineering for data generation.
  • Familiarity with dataset distillation and bias mitigation.