Pylon
Prompt evaluation
Prompt iterations rubber-banded: each engineer's fix overcorrected the last. Evals are now a merge requirement—no eval, no commit.
- 10 mins of active debugging eliminated per on-call incident
Teams couldn't manually review hundreds of daily AI hotel calls. Audio models now evaluate raw recordings, routing exceptions to humans.
An all-in-one travel and expense management platform with an AI voice agent making hundreds of calls daily to hotels worldwide to confirm bookings and deliver payment card details on behalf of travelers.
As call volume scaled to hundreds per day, the development and ops teams couldn't manually listen to each conversation to assess quality or...
“When we started to build this AI agent very quickly, what started happening was that hundreds of calls started happening on behalf of our travelers.”
All-in-one travel and expense management platform that combines corporate travel booking, expense tracking, and corporate card services.
AI observability and evaluation platform that helps developers build, test, and monitor LLM-powered applications.
Navan's Voice call evaluation is part of this use case:
Related implementations across industries and use cases
Prompt iterations rubber-banded: each engineer's fix overcorrected the last. Evals are now a merge requirement—no eval, no commit.
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
Reviewers struggled to predict how code ripples through the system. AI now flags cross-service risks that cause outages.
Mining 5 coachable minutes from a 45-min call kept coaching rare. AI summaries in 20 seconds lifted scorecards from 23 to 218 in a quarter.
Student insights were trapped in unreviewed audio. AI securely evaluates every call to power instant feedback and proactive coaching.
Scattered data and basic coding tools bottlenecked engineers. A 9-agent AI workflow shifts them from writing code to directing AI teams.
Sequential AI testing bottlenecked development. Engineers built a concurrent, code-first pipeline to evaluate agent responses in seconds.
Tournaments running simultaneously meant an hour of manual checks each. AI agents now run them in minutes, freeing the team to be proactive.
Large AI training jobs meant fighting for preemptible slots or leaving campus. Marlowe gave any lab guaranteed multi-node access on demand.
Teams couldn't manually review hundreds of daily AI hotel calls. Audio models now evaluate raw recordings, routing exceptions to humans.
An all-in-one travel and expense management platform with an AI voice agent making hundreds of calls daily to hotels worldwide to confirm bookings and deliver payment card details on behalf of travelers.
As call volume scaled to hundreds per day, the development and ops teams couldn't manually listen to each conversation to assess quality or...
“When we started to build this AI agent very quickly, what started happening was that hundreds of calls started happening on behalf of our travelers.”
All-in-one travel and expense management platform that combines corporate travel booking, expense tracking, and corporate card services.
AI observability and evaluation platform that helps developers build, test, and monitor LLM-powered applications.
Navan's Voice call evaluation is part of this use case:
Related implementations across industries and use cases
Prompt iterations rubber-banded: each engineer's fix overcorrected the last. Evals are now a merge requirement—no eval, no commit.
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
Reviewers struggled to predict how code ripples through the system. AI now flags cross-service risks that cause outages.
Mining 5 coachable minutes from a 45-min call kept coaching rare. AI summaries in 20 seconds lifted scorecards from 23 to 218 in a quarter.
Student insights were trapped in unreviewed audio. AI securely evaluates every call to power instant feedback and proactive coaching.
Scattered data and basic coding tools bottlenecked engineers. A 9-agent AI workflow shifts them from writing code to directing AI teams.
Sequential AI testing bottlenecked development. Engineers built a concurrent, code-first pipeline to evaluate agent responses in seconds.
Tournaments running simultaneously meant an hour of manual checks each. AI agents now run them in minutes, freeing the team to be proactive.
Large AI training jobs meant fighting for preemptible slots or leaving campus. Marlowe gave any lab guaranteed multi-node access on demand.