Job Summary:
We are seeking an experienced Agentic AI Evaluation Data Scientist to design and implement evaluation frameworks for conversational and autonomous AI systems supporting commerce and customer-facing use cases. The role will be responsible for assessing agent performance across task completion, tool-use accuracy, multi-turn coherence, recommendation relevance, hallucination, and customer experience outcomes. The successful candidate will build automated and human-in-the-loop evaluation pipelines using Azure AI Foundry, establish regression and red-teaming practices, and analyze production telemetry to identify quality drift, failure patterns, bias, and fairness concerns. Working closely with AI Engineers and Solution Architects, the role will embed evaluation gates into MLOps and CI/CD pipelines and provide clear risk assessments to support responsible AI governance and release decisions.
Responsibilities:
- Design and implement evaluation frameworks for agentic AI systems (conversational shopping assistants, autonomous task agents), covering task success rate, tool-call accuracy, and multi-turn coherence.
- Build offline and online evaluation pipelines using Azure AI Foundry evaluation tools, combining LLM-as-judge techniques, human-in-the-loop review, and statistical significance testing.
- Define and track quality metrics specific to commerce agents — recommendation relevance, hallucination rate, task completion, escalation/handoff accuracy, and customer satisfaction proxies (CSAT, resolution rate).
- Design red-teaming and adversarial testing protocols to probe agent robustness against prompt injection, jailbreaks, and edge-case commerce scenarios (returns, fraud, pricing errors).
- Develop regression testing suites for agent to detect quality drift across model versions, prompt changes, and RAG knowledge base updates.
- Analyze production agent logs/telemetry to identify failure patterns, bias, and fairness issues, feeding insights back into prompt engineering, fine-tuning, and RAG retrieval improvements.
- Collaborate with AI Engineers and Solution Architects to instrument agents for observability (tracing, logging via Azure Monitor/Application Insights) and define evaluation gates within MLOps/CI-CD pipelines.
- Communicate evaluation findings and risk assessments to technical and business stakeholders, informing go/no-go decisions for agent releases and supporting responsible AI governance/compliance reporting.
Preferred Certifications:
- Microsoft Certified: Azure AI Engineer Associate (AI-102)
- Microsoft Certified: Azure Data Scientist Associate (DP-100)
- Microsoft Certified: Azure Fundamentals / AI Fundamentals (AZ-900 / AI-900) — baseline
- Relevant Responsible AI / LLM evaluation credentials (e.g., DeepLearning.AI courses on LLM evaluation, Azure AI Content Safety specialization)
- Statistics/ML academic background (MS/PhD in relevant field) — valued though not mandatory given strong practical evaluation experience
Skills Required:
- 5–8 years in data science/ML, with at least 2–3 years specifically focused on evaluating LLM-based or agentic AI systems in production, ideally within retail, e-commerce, or customer-facing conversational AI domains.
Benefits:
Joining Cognizant will give you the opportunity to learn and collaborate with some of the most talented people in the industry, while having your finger on the pulse of emerging industry trends and working on the cutting edge of technology in your field of expertise. We recognize that our people perform at their best when they feel valued as significant contributors and that is why at Cognizant, taking care of our employees is a priority:
- You can pursue innovative career tracks and opportunities here
- You can enhance your professional development through education and dedicated training
- We'll give you the skills you need to keep pace with the changing workplace while our compensation, benefits and wellness packages help you stay healthy and plan for the future.