Skip to main content
POST
Create Custom Metric
Integration Prompt for AI Agents
This endpoint allows you to create a custom metric for your agent. Provide the necessary details to add a metric to your agent.

Headers

X-Organization-Id
string
X-API-Key
string
required

API key required to authenticate requests.

Body

application/json

Request model for creating a custom metric

name
string
required

Name of the custom metric

response_type
enum<string>
required

Type of response expected

Available options:
pass_fail,
yes_no,
qualitative,
quantitative,
json,
enum,
int,
float,
boolean,
tool_call
bluejay_as_code_id
string<uuid> | null

Stable code-addressable identifier

agent_id
integer | null
deprecated

ID of the agent this metric belongs to (deprecated, use agent_ids instead)

agent_ids
integer[] | null

List of agent IDs to associate this metric with

prompt
string | null

LLM-judge system prompt (required for llm_judge; null for other types)

summary
string | null

Short plain-English description of the metric (LLM-generated)

metric_type
string
default:llm_judge

Specific metric kind (llm_judge, tool_call, audio_*, vocal_fry); legacy alias: type

settings
Settings · object

Type-specific config (e.g. tool_call: {tool_names: [...]})

eval_segment
enum<string>
default:all

Which side of the conversation is evaluated

Available options:
agent,
user,
all
scope
enum<string>
default:all

Which pipeline runs this metric

Available options:
observability,
simulations,
all
min_value
number | null

Minimum value for quantitative metrics

max_value
number | null

Maximum value for quantitative metrics

category
string | null

Category for organizing metrics

tags
string[] | null

Tags for categorizing the metric

scoring_guidance
string | null

Guidance on how to score this metric

eval_modality
enum<string> | null
default:AUTO

Evaluation modality AUTO/TEXT/AUDIO; legacy alias: eval_route

Available options:
AUDIO,
TEXT,
AUTO
model
string | null

Model to use for text evaluation of the metric

audio_model
string | null

Model to use for audio evaluation of the metric

temperature
number | null

Temperature for evaluating the metric

enum_options
string[] | null

Options for enum metrics

tool_names
string[] | null

Expected tool names for tool_call metrics

json_schema
Json Schema · object | null

JSON Schema describing the object the judge must return (required for json metrics)

allow_not_applicable
boolean
default:false

Whether this metric allows 'not applicable' responses.

template_id
string<uuid> | null

Library template this metric was created from

Response

Successful Response

Response model for custom metric operations

id
string
required

Unique identifier for the custom metric

name
string
required

Name of the custom metric

response_type
enum<string>
required

Type of response expected

Available options:
pass_fail,
yes_no,
qualitative,
quantitative,
json,
enum,
int,
float,
boolean,
tool_call
created_at
string<date-time>
required

When this metric was created

bluejay_as_code_id
string | null

Stable code-addressable identifier (same as id for custom metrics)

prompt
string | null

LLM-judge system prompt for this metric (null for non-judge types)

description
string | null

Deprecated: same value as prompt

summary
string | null

Short plain-English description of the metric (LLM-generated)

metric_type
string
default:llm_judge

Specific metric kind

eval_method
string
default:llm_judge

Evaluation category (llm_judge|statistical|ml_model|deterministic)

settings
Settings · object

Type-specific config

eval_segment
string
default:all

Which side of the conversation is evaluated (agent|user|all)

scope
string
default:all

Which pipeline runs this metric (observability|simulations|all)

agent_ids
integer[]

List of agent IDs this metric belongs to

min_value
number | null

Minimum value for quantitative metrics

max_value
number | null

Maximum value for quantitative metrics

category
string | null

Category for organizing metrics

tags
string[] | null

Tags for categorizing the metric

scoring_guidance
string | null

Guidance on how to score this metric

updated_at
string<date-time> | null

When this metric was last updated

updated_by
string | null

User who last updated this metric

created_by
string | null

User who created this metric

eval_modality
enum<string> | null

Evaluation modality AUTO/TEXT/AUDIO

Available options:
AUDIO,
TEXT,
AUTO
model
string | null

Model to use for text evaluation of the metric

audio_model
string | null

Model to use for audio evaluation of the metric

temperature
number | null

Temperature for evaluating the metric

enum_options
string[] | null

Options for enum metrics

tool_names
string[] | null

Deprecated: mirror of settings.tool_names for tool_call metrics

json_schema
Json Schema · object | null

JSON Schema describing the object the judge must return (json metrics only)

allow_not_applicable
boolean
default:false

Whether this metric allows 'not applicable' responses.

eval_modality_auto_selected
boolean
default:false

Deprecated: always false (column removed).

template_id
string | null

Library template this metric was created from