Navan
Voice call evaluation
Teams couldn't manually review hundreds of daily AI hotel calls. Audio models now evaluate raw recordings, routing exceptions to humans.
- Over 0.9 macro F1 across all eval control groups
- Eval system F1 improved from 0.56 to 0.89