Senior Research Scientist, Agentic AI
Job Description
We build the evaluation layer that validates AI agents before they reach customers — an automated system that scores across large volumes of agent traces.
You'll own core parts of that platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on.
What you'll do
• Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety — across LLM-as-a-Judge, trajectory-based, and human evaluation
• Take problems from research question to prototype to shipped feature, owning them end to end
• Build and harden the pipelines and scoring logic behind customer-facing evaluation
• Curate synthetic and real-world datasets; measure the evaluator itself for consistency and agreement with human labels
Requirements
Department: Engineering, Infrastructure and Operations
Function: Engineering
Experience Level: Not Applicable