Dropbox
Model evaluation
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
- Under 10 minutes for automated PR evaluations
Engineering handoffs delayed prompt updates. Now, behavioral researchers test and deploy conversation changes directly to production.
A developer of an AI companion app designed to build authentic, non-romantic relationships through natural voice conversations and complex memory systems.
Evaluating subjective nuances like emotional intelligence, conversational pacing, and natural memory recall proved impossible using automated metrics...
“How will humanity build healthy relationships with AI? That's the question Tolan explores.”
AI-powered virtual companion app for emotional support and conversation.
AI observability and evaluation platform that helps developers build, test, and monitor LLM-powered applications.
Portola's Prompt evaluation is part of this use case:
Related implementations across industries and use cases
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
Teams couldn't manually review hundreds of daily AI hotel calls. Audio models now evaluate raw recordings, routing exceptions to humans.
Manual review of sensitive files took two days. AI agents now finish the work in one hour.
800-student courses, limited TAs. AI grader clusters mistakes for instructor review; students get 24/7 tutoring that insists they think first.
23% of turns were disengaging — Cait only found out after each conversation ended. State inference now runs live and adjusts mid-turn.
Scattered data and basic coding tools bottlenecked engineers. A 9-agent AI workflow shifts them from writing code to directing AI teams.
Sequential AI testing bottlenecked development. Engineers built a concurrent, code-first pipeline to evaluate agent responses in seconds.
On-premise systems, dispersed and brittle, bottlenecked every release. AI agents now run routine dev steps — hours cut to minutes.
A mistranslated word could derail global R&D projects. Now, researchers instantly refine technical papers & communicate seamlessly across languages.
Engineering handoffs delayed prompt updates. Now, behavioral researchers test and deploy conversation changes directly to production.
A developer of an AI companion app designed to build authentic, non-romantic relationships through natural voice conversations and complex memory systems.
Evaluating subjective nuances like emotional intelligence, conversational pacing, and natural memory recall proved impossible using automated metrics...
“How will humanity build healthy relationships with AI? That's the question Tolan explores.”
AI-powered virtual companion app for emotional support and conversation.
AI observability and evaluation platform that helps developers build, test, and monitor LLM-powered applications.
Portola's Prompt evaluation is part of this use case:
Related implementations across industries and use cases
Scattered spreadsheets couldn't catch AI hallucinations. Now, automated LLM judges evaluate every prompt change to block regressions.
Teams couldn't manually review hundreds of daily AI hotel calls. Audio models now evaluate raw recordings, routing exceptions to humans.
Manual review of sensitive files took two days. AI agents now finish the work in one hour.
800-student courses, limited TAs. AI grader clusters mistakes for instructor review; students get 24/7 tutoring that insists they think first.
23% of turns were disengaging — Cait only found out after each conversation ended. State inference now runs live and adjusts mid-turn.
Scattered data and basic coding tools bottlenecked engineers. A 9-agent AI workflow shifts them from writing code to directing AI teams.
Sequential AI testing bottlenecked development. Engineers built a concurrent, code-first pipeline to evaluate agent responses in seconds.
On-premise systems, dispersed and brittle, bottlenecked every release. AI agents now run routine dev steps — hours cut to minutes.
A mistranslated word could derail global R&D projects. Now, researchers instantly refine technical papers & communicate seamlessly across languages.