What You’ll Learn
- What Custom Metrics are and why they matter
- How to define evaluation criteria for your specific use case
- How Custom Metrics integrate with simulations and observability
How Custom Metrics Work
You create Custom Metrics to score conversations on domain-specific behavior such as compliance, resolution quality, empathy, or escalation accuracy. Those metrics can then be reused across simulations and production evaluations. Custom Metrics support two evaluation modes: LLM-as-a-Judges that use a prompt to score conversations, and formula-based definitions that compute composite scores from other metrics. You can prototype and refine metrics in Metrics Lab before deploying them.Key Capabilities
Click each capability to see what it gives you.LLM-as-a-Judge
LLM-as-a-Judge
Write a natural-language prompt that scores conversations on any criteria you define. The judge runs against the call transcript and returns a verdict shaped by your prompt and the metric’s response type.
Formula metrics
Formula metrics
Combine existing metric scores using arithmetic expressions to create composite indicators. Useful for “call quality” scores that weight several signals together.
Cross-workflow reuse
Cross-workflow reuse
The same metric works in both simulation and observability evaluations. Define it once and apply it to test runs and production calls.
Metrics Lab integration
Metrics Lab integration
Test scoring logic against sample transcripts in Metrics Lab before pointing the metric at production. Iterate on the prompt until the judge agrees with your spot-checks.
Common Use Cases
- Score whether an agent correctly verified a customer’s identity before sharing account details
- Track empathy and de-escalation quality across production calls
- Create a composite “call quality” score that weights resolution, tone, and compliance together
Resources
Metric Types
Pick the right response type for each measurement goal.
Prompting Guide
Write LLM-as-a-Judge prompts that score consistently across runs.
Dynamic Variables
Inject call-specific context into metric definitions at evaluation time.
Metrics Lab
Prototype and test metrics against sample transcripts before deployment.
Create Custom Metric API
Define a new Custom Metric programmatically.
Evaluate Endpoint
Submit calls for evaluation and pass metadata for dynamic variable substitution.