Staff Test Engineer - AI
Outreach · 78 days ago
About Outreach
Outreach, founded in 2014, is the only complete agentic AI [ platform for revenue teams. Outreach infuses agentic AI, conversation intelligence, and assistive AI to power hundreds of use cases across revenue motions. From new logo prospecting to expansions, deal acceleration, driving retention, and forecasting, Outreach AI automates workflows and frees sellers to focus on more strategic conversations and actions. Revenue leaders benefit from connected account visibility, performance insights, and higher forecasting accuracy across every GTM team. World leading enterprise organizations use Outreach to power their revenue teams, including Databricks, SAP, Siemens, and Verizon to name a few.
\n
Your Daily Adventures Will Include: Responsibilities:
- Own the AI Quality Strategy. Define and lead the end-to-end testing strategy for Outreach’s GenAI platform, including agentic workflows, LLM tool calls, LangGraph orchestration, and supporting ML pipelines.
- Build Evaluation Frameworks. Design and implement evaluation systems that handle both deterministic and non-deterministic outputs — combining rule-based assertions, golden dataset testing, and LLM-as-Judge approaches to grade agent responses at scale.
- Test Agents End-to-End. Own testing across Outreach’s suite of AI agents — Revenue Agent, Research Agent, Meeting Agent, Personalisation Agent, and Ask Outreach — covering functional correctness, tool selection accuracy, context handling, and response quality.
- Partner with DS and Engineering. Work closely with Data Science, MLOps, and platform engineers to ensure testability is designed in from the start — not bolted on after.
- Drive CI/CD for AI. Integrate evaluation pipelines into CI/CD workflows so that regressions in agent behavior are caught before they reach production.
- Define Quality Metrics. Establish and track metrics that matter for AI systems: answer quality scores, tool invocation accuracy, hallucination rates, latency, and regression trends over model and prompt changes.
- Champion Best Practices. Define standards for AI testing across the org — including prompt regression testing, retrieval quality evaluation, and agent behavior contracts.
- Mentor and Influence. Raise the quality bar across engineering teams by mentoring engineers, reviewing designs for testability, and advocating for quality-driven development practices.
- Stay Current. Actively track developments in AI evaluation tooling, LLM benchmarking, and testing research — and bring relevant advances into our practice.
Our Vision of You: Minimum Qualifications:
- 7–12 years of experience in software development and/or test automation, with demonstrated experience leading quality efforts on complex, distributed systems.
- B.S. in Computer Science or a related technical field.
- Strong programming skills in Python, with experience writing reusable, maintainable test frameworks.
- Proven experience testing large-scale backend or platform systems, including microservices and API layers.
- Deep understanding of test design principles, CI/CD integration, and scalable test automation.
- Experience with test frameworks such as PyTest or equivalent.
- Solid understanding of evaluation methodologies for non-deterministic systems — including statistical assertions, behavioral testing, and regression baselines.
- Hands-on experience with Databricks for building and validating ML pipelines and data workflows.
- Experience with MLflow for experiment tracking, model versioning, and pipeline observability.
- Strong communication and collaboration skills across engineering, data science, and product functions.
Preferred Qualifications:
- Experience testing GenAI products, LLM-based systems, or agentic AI platforms.
- Experience with prompt engineering and prompt tuning — understanding how prompt changes affect model behavior and building regression suites to catch prompt-driven regressions.
- Hands-on experience with LLM-as-Judge evaluation patterns — using LLMs to grade LLM outputs at scale.
- Familiarity with LangGraph, LangChain, or similar agent orchestration frameworks.
- Experience with ML pipelines, ML flow tooling (e.g., MLflow, Kubeflow, Metaflow), or model evaluation workflows.
- Understanding of RAG (Retrieval-Augmented Generation) architectures and how to evaluate retrieval quality.
- Experience with cloud platforms (AWS, GCP, or Azure) and containerized environments (Docker, Kubernetes).
- Domain knowledge in sales, sales engagement, or CRM platforms (e.g., Salesforce, HubSpot, or similar) — understanding the workflows, terminology, and data that sales teams operate with.
- Prior experience contributing to AI quality strategies in a product or research environment.
\n
Why You’ll Love It Here
● Highly competitive salary
● 25 days annual vacation time + sick time and casual leave
● Group medical policy coverage available to employees and up to 5 eligible family members
● OPD benefit covered up to INR 10,000
● Life insurance and personal accident insurance at 3x annual CTC
● 26 weeks of maternity leave pay, and 15 days of paternity leave pay
● Opportunity to be part of company success via the RSU program
● Diversity and inclusion programs that promote employee resource groups like OWN+ (Outreach Women's Network), Adelante (Latinx community), OBX (Outreach Black Connection), Mosaic (AAPI community), Pride (LGBTQIA+), Gender+, Disability Community, and Veterans/Military
● Employee referral bonuses to encourage the addition of great new people to the team
● Fun company and team outings because we play just as hard as we work
Outreach is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Our success is reliant on building teams that include people from different backgrounds and experiences who can elevate assumptions and ideas with fresh perspectives. We're dedicated to hiring the whole human, not just a resume. To that end, we look for a diverse pool of applicants-including those from historically marginalized groups. We would like to invite you to apply even if you don't think you meet all of the requirements listed below. We don't want a few lines in a job description to get between us and the opportunity to meet you.
