Dropbox
Model evaluation
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
- Under 10 minutes for automated PR evaluations
Teams couldn't manually review hundreds of daily AI hotel calls. Audio models now evaluate raw recordings, routing exceptions to humans.
An all-in-one travel and expense management platform with an AI voice agent making hundreds of calls daily to hotels worldwide to confirm bookings and deliver payment card details on behalf of travelers.
As call volume scaled to hundreds per day, the development and ops teams couldn't manually listen to each conversation to assess quality or...
“When we started to build this AI agent very quickly, what started happening was that hundreds of calls started happening on behalf of our travelers.”
All-in-one travel and expense management platform that combines corporate travel booking, expense tracking, and corporate card services.
AI observability and evaluation platform that helps developers build, test, and monitor LLM-powered applications.
Navan's Voice call evaluation is part of this use case:
Related implementations across industries and use cases
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
Reviewers struggled to predict how code ripples through the system. AI now flags cross-service risks that cause outages.
Dense API docs bottlenecked non-technical users. Now, Claude builds fully configured voice agents from simple plain-text requests.
Mining 5 coachable minutes from a 45-min call kept coaching rare. AI summaries in 20 seconds lifted scorecards from 23 to 218 in a quarter.
Student insights were trapped in unreviewed audio. AI securely evaluates every call to power instant feedback and proactive coaching.
Scattered data and basic coding tools bottlenecked engineers. A 9-agent AI workflow shifts them from writing code to directing AI teams.
Sequential AI testing bottlenecked development. Engineers built a concurrent, code-first pipeline to evaluate agent responses in seconds.
On-premise systems, dispersed and brittle, bottlenecked every release. AI agents now run routine dev steps — hours cut to minutes.
A mistranslated word could derail global R&D projects. Now, researchers instantly refine technical papers & communicate seamlessly across languages.
Teams couldn't manually review hundreds of daily AI hotel calls. Audio models now evaluate raw recordings, routing exceptions to humans.
An all-in-one travel and expense management platform with an AI voice agent making hundreds of calls daily to hotels worldwide to confirm bookings and deliver payment card details on behalf of travelers.
As call volume scaled to hundreds per day, the development and ops teams couldn't manually listen to each conversation to assess quality or...
“When we started to build this AI agent very quickly, what started happening was that hundreds of calls started happening on behalf of our travelers.”
All-in-one travel and expense management platform that combines corporate travel booking, expense tracking, and corporate card services.
AI observability and evaluation platform that helps developers build, test, and monitor LLM-powered applications.
Navan's Voice call evaluation is part of this use case:
Related implementations across industries and use cases
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
Reviewers struggled to predict how code ripples through the system. AI now flags cross-service risks that cause outages.
Dense API docs bottlenecked non-technical users. Now, Claude builds fully configured voice agents from simple plain-text requests.
Mining 5 coachable minutes from a 45-min call kept coaching rare. AI summaries in 20 seconds lifted scorecards from 23 to 218 in a quarter.
Student insights were trapped in unreviewed audio. AI securely evaluates every call to power instant feedback and proactive coaching.
Scattered data and basic coding tools bottlenecked engineers. A 9-agent AI workflow shifts them from writing code to directing AI teams.
Sequential AI testing bottlenecked development. Engineers built a concurrent, code-first pipeline to evaluate agent responses in seconds.
On-premise systems, dispersed and brittle, bottlenecked every release. AI agents now run routine dev steps — hours cut to minutes.
A mistranslated word could derail global R&D projects. Now, researchers instantly refine technical papers & communicate seamlessly across languages.