Skip to main content
Every Custom Metric has a response_type that determines the shape of the score Bluejay produces. Choose the type that matches how you want to reason about the result downstream, whether that’s in dashboards, alerts, or workflows.

Pass / Fail

response_type: pass_fail
Use Pass / Fail for binary compliance checks, required steps, or any criteria where the outcome is clear-cut. Results aggregate cleanly into pass rates on your dashboard.
The most common type. Bluejay returns either pass or fail based on whether the conversation meets your criteria. To create this metric via API, use the Create Custom Metric endpoint with response_type set to pass_fail. The endpoint accepts optional fields such as agent_id, agent_ids, category, tags, and allow_not_applicable. See the full request schema on that page.

Quantitative

response_type: quantitative
Use Quantitative when gradations matter, such as tone of voice, resolution depth, or explanation clarity. Set explicit min_value and max_value so Bluejay knows the scale boundaries.
Bluejay returns a numeric score within the range you define using min_value and max_value. Use this for nuanced scoring on a scale. To create this metric via API, use the Create Custom Metric endpoint with response_type set to quantitative and include min_value and max_value so Bluejay knows the allowed range.

Qualitative

response_type: qualitative
Use Qualitative for feedback that needs nuance and context, such as coaching notes, tone summaries, or structured observations. Because results are text, they don’t aggregate numerically, so pair them with a quantitative or pass/fail metric if you also need trend tracking.
Bluejay returns a free-form text summary describing its assessment. Useful for generating narrative feedback rather than a structured score. To create this metric via API, use the Create Custom Metric endpoint with response_type set to qualitative. No extra type-specific fields are required beyond name and description.

Enum

response_type: enum
Use Enum when you need consistent, controlled labels, such as call outcomes, escalation reasons, or issue categories. The discrete labels make it easy to group and filter results across many conversations.
Bluejay classifies the conversation into one of the exact labels you define in enum_options. This is ideal for categorization tasks. To create this metric via API, use the Create Custom Metric endpoint with response_type set to enum and include enum_options as an array of allowed labels.

JSON

response_type: json
Use JSON when you want to extract multiple structured signals from a single evaluation. Keep the output schema well-defined in your description so Bluejay consistently produces the same shape.
Bluejay returns a structured JSON object, letting you extract multiple signals from a single metric evaluation in one pass. To create this metric via API, use the Create Custom Metric endpoint with response_type set to json. Describe the desired object shape clearly in description so evaluations stay consistent.

Tool Call

response_type: tool_call A deterministic check that passes when the agent called the tool(s) you name. No LLM is involved: Bluejay compares the tool_names you configure against the tool calls captured for the conversation.
Use Tool Call to verify required actions actually happened, like “the agent must call create_ticket on every support call.” Pair it with a native provider integration (Vapi, Retell, ElevenLabs, LiveKit) so tool calls are captured automatically.
To create this metric via API, use the Create Custom Metric endpoint with response_type set to tool_call and the expected names in tool_names.

Yes / No (Deprecated)

response_type: yes_no
Yes / No is deprecated. Use Pass / Fail instead. It is functionally identical and is the preferred type going forward.
Semantically identical to Pass / Fail, but the framing uses question-style prompts. Bluejay returns yes or no. To create this metric via API, use the Create Custom Metric endpoint with response_type set to yes_no. Optional fields such as agent_id, category, and tags match other metric types. See the full schema on that page.

Quick Reference


Not Applicable

Any metric type can be configured with allow_not_applicable: true. When enabled, Bluejay may return Not Applicable if the criteria simply doesn’t apply to a given conversation. For example, a transfer metric on a call that never needed a transfer.
Set allow_not_applicable when creating or updating the metric via the Create Custom Metric or Update Custom Metric endpoints.

Resources

Prompting Guide

Write LLM-as-a-Judge prompts that score consistently for any response type.

Dynamic Variables

Inject call-specific context into your metric definitions at evaluation time.

Metrics Lab

Prototype and refine metrics against sample transcripts before going live.

Create Custom Metric API

Define a new Custom Metric programmatically with any response type.