Skip to main content
ImprovementAPI

Agent versions

New: agent versions live at /v1/agents/{id}/versions (create, list, get, delete, promote, labels).Same as before: POST /v1/evaluate and /evaluate-call. New optional integer version. label still works. Omit version and we use the version that currently has the production label. An unknown version uses the production version and returns a warning. An unknown label is created and attached to the production version.Deprecated: /v1/agents/{id}/prompts paths still work. Get-by-label stays on GET /prompts?label= (no /versions twin).New: create/queue simulation accepts agent_version_id.Deprecated: prompt_id on create/queue simulation. Same uuid, still works. If both are sent and they differ, the request is 400.Evaluate · Create Agent Version
NewIntegrations

Better Stack paging

Alerts and uptime monitors can now page through Better Stack alongside PagerDuty. Paste a Better Stack Uptime API token under Settings → Integrations, then switch on Better Stack Notifications on any alert or monitor. DOWN opens an incident and recovery resolves it, with nothing to configure on the Better Stack side.Better Stack integration · PagerDuty integration
NewObservability

Issues

Production failures now group into Issues you can triage. Patterns appear once the same problem recurs across calls, with a title. Known errors (a tool 500, a tool that should have run but never did) appear as soon as they are recognized. Open Monitor → Issues for the full list, grouped by status. The list is not paginated.Issues · Monitor → Issues · List Issues API
NewObservability

Topics (early access)

Bluejay now clusters what’s happening across your production calls automatically. The Topics map is an interactive view of what your callers talked about: pan, zoom, search, and see trends against the prior window. Rolling out to select organizations first. Contact us for access.Observability overview
ImprovementDigital Humans

Expanded language coverage

Digital Humans picked up a wave of new languages and dialects this summer: Urdu, Hebrew, Tagalog, Telugu, Tamil, Malayalam, Kannada, Marathi, and Gujarati, plus Egyptian and Levantine Arabic with native dialect voices. Vietnamese and Cantonese voice coverage was also filled out, and every non-English language received dialect-aware prompting for more authentic speech.Digital Humans overview
NewSimulations

Custom SIP headers

Simulations can now attach custom SIP headers to the calls they place. Headers configured on the agent are delivered on the SIP INVITE, so agents behind infrastructure that routes or authenticates on custom headers can be tested exactly as they run in production.SIP connections
NewMetrics

Custom metrics judged on call audio

Audio custom metrics are now judged natively on the call recording by an audio-capable model, with optional knowledge-base grounding. Tone, pacing, pronunciation, and other things a transcript can’t capture are now first-class evaluation targets.Custom metrics overview
NewMetrics

Outbound call-outcome metrics & spoken-phrase checks

Eight new library metrics for outbound calling: LLM-judged outcomes (human picked up, voicemail detected, transfer attempted, stuck call, abrupt termination, natural delivery) plus deterministic required-disclosure and prohibited-wording checks. Phrase metrics are authored by simply typing the words the agent must (or must not) say — Bluejay compiles the matching and absorbs ASR variance like punctuation and acronym spellings automatically.Metric types
NewIntegration

Twilio account linking

Link your own Twilio account from Settings → Integrations, so simulations can dial using your Twilio numbers.
ImprovementIntegration

Pipecat improvements

Bluejay’s Pipecat Cloud integration now sends the run’s metadata in the start-call body, so your Pipecat agent can identify and branch on the Bluejay test run that’s calling it. Sessions also end cleanly when the agent hangs up or fails to start, instead of waiting for a timeout.Pipecat simulations
ImprovementBilling

Spend breakdown in credits

The Spend page now breaks your usage down per product, denominated in credits, and the usage page breaks credit consumption down by agent and by individual simulation — including observability-only agents.
NewIntegration

Bland integration

Bland is now a fully supported provider. Connect from Create Agent with inline API key setup, pick a Conversational Pathway, and pull, edit, push, and publish pathway versions from an interactive canvas inside Bluejay. Simulations run over Bland’s native voice bridge (with PSTN dialing also supported) and record the exact path each call took through the pathway graph, and Bland production calls can be ingested into observability.Bland Observability · Bland Simulations
NewObservability

Natural-language log filtering

Filter your call logs by just typing. Plain-English queries are translated into structured log filters automatically, on top of the freeform full-text search over call summaries that shipped earlier in July.Observability overview
ImprovementPlatform

Knowledge bases: URL & HTML ingestion

Knowledge bases now accept web pages directly by URL — Bluejay crawls and ingests the page server-side — as well as uploaded HTML files, alongside the existing document upload flow.
NewIntegration

Amelia integration

Test Amelia agents on Bluejay over both chat and voice. Enter your Amelia credentials in Settings → Integrations, pick a domain and flow, and sync agents into Bluejay. You can view and edit the agent’s flow on an interactive canvas and push changes back to Amelia.Amelia integration
NewReports

Deep Reports

Agent-generated, verified reports built from your simulation and observability data. A report studio lets you edit the report through chat on a Bluejay-native canvas, schedule recurring sends, and email the result — with PDF export for run comparisons.
NewDashboards

Dashboard widgets: histograms, filters, and 32 new metrics

A major dashboard upgrade: a histogram widget for numeric metric distributions, 32 newly graphable metrics in the widget dialogs, distribution aggregation for integer-scale custom metrics, per-widget conversation filters, metadata Group By, and click-a-threshold-bar to filter the explorer to that group. The Logs page also gained a Table | Chart toggle to chart whatever view you’ve filtered to.Dashboards
NewMetrics

Answer Latency

Every simulated call now measures how long the agent took to answer, across all transports. Answer Latency shows up in the conversation Evals tab and as a sortable column in simulation run results, so slow pickup is a trackable signal.
NewMetrics

Human review overrides

Reviewers can override an automated metric verdict by hand. The human correction is stored alongside the original AI grade — both are preserved — and corrected values show up everywhere: tables, aggregates, and reports.Custom metrics overview
NewIntegration

Bluejay in Slack

Ask Bluejay about your agents, simulations, and production conversations directly from Slack. The per-org Slack bot supports threaded follow-ups without re-tagging, works in externally shared channels, and performs write actions safely via one-click approval messages.Slack integration
NewSimulations

Outbound trigger webhook

For agents that dial out from your own system, configure a trigger webhook in agent settings: Bluejay calls it to tell your infrastructure to place the call (or send the SMS) to the Digital Human. Works with any provider, so outbound agents are testable without Bluejay controlling the dialer.
NewSimulations

Extension dialing (DTMF)

Digital Humans can now dial a configured extension as DTMF tones after connecting, so agents behind IVR menus and extension routing can be tested end to end. Extensions can be set per Digital Human (including in bulk) and on red-team launches.
ImprovementIntegration

Google Cloud: selective sync & keyless setup

Choose exactly which Google Conversational Agents to sync (or toggle sync-all), and connect your Google Cloud project without service-account keys using least-privilege Workload Identity Federation, configured from the settings UI with a paste-and-run Cloud Shell flow.Keyless Google Cloud setup
ImprovementMetrics

Grounded, customizable judging

LLM-judge custom metrics can now ground their evaluation in your knowledge base, so factual-accuracy metrics grade against your actual documentation. Standard hallucination and redundancy metrics accept per-org prompt overrides and are editable like custom metrics, and new default eval models improve text and audio judging quality.Prompting guide
ImprovementDigital Humans

Digital Human management upgrades

A table view of Digital Humans with inline per-field editing, per-Digital-Human run history pages, duplicate-to-a-target-count, and two new generation sources: build Digital Humans from your knowledge base’s topics, or derive traits from your agent’s template variables so provider call variables are filled with each persona’s data.Digital Humans overview
NewObservability

Trace view v2 with token costs

A redesigned trace experience: per-leg latency dashboards, a rebuilt trace side panel with a standalone full view, and trace export alongside simulation results. Model spans now show token usage (input/cached/output, text/audio splits) with estimated cost per model and trace-level totals. ElevenLabs agent traces are ingested via OpenTelemetry, so their tool calls and spans appear like any other provider’s.Traces
NewPlatform

Self-serve signup

Bluejay is open to self-serve signups. Sign up with email verification, land in a redesigned onboarding flow with Bluejay AI as your first experience, and start testing immediately against seeded sample agents — including read-only “Bug Bounty” challenge agents built to be probed.Quickstart
NewMetrics

Metrics Library

An app-store-style library at /metrics: browse standard and community-published metric templates, preview them, rate them, and add them to your org in one click. You can publish your own metrics to the library through a review queue, and My Metrics was restructured to match.Custom metrics overview
NewSimulations

Red teaming revamp

A rebuilt red-team experience: a dedicated launch flow and results page with an OWASP/ATLAS-mapped scorecard and leak table, opt-in content-safety attacks as a distinct attack category, audio-native attacker behaviors, keypad (DTMF) attacks, and extension support for targets behind IVRs.Red teaming · Content safety
NewSimulations

Import workflow diagrams & IVR test plans

Drop a workflow diagram (for example an IVR flowchart image) into Bluejay AI and it becomes an editable Agent Workflow with a generated checklist-style test plan covering the flow’s paths. Step-based IVR test-plan exports import too: Bluejay turns them into transcript-replay Digital Humans, so existing IVR test scripts become runnable simulations.Agent workflows
NewMetrics

Audio Clarity Score

A single 0–100 audio clarity score on every call, replacing the earlier burst-SNR approach. The agent channel is scored for intelligibility and the caller channel with DNSMOS, calibrated against a human-labeled gold set, so brief bad stretches aren’t averaged away.
NewSimulations

Multimodal journeys & follow-up SMS

Customer journeys can now mix voice and SMS — for example an SMS leg that ends when the call leg starts — with per-step success criteria and transition guards. Follow-up SMS gets full configuration: call association, a sender allowlist on the Phone Numbers page, and a per-simulation SMS inactivity timeout.Customer journeys · Follow-up SMS
ImprovementSimulations

Smarter call endings

Whether a simulated call should end is now decided by a consensus of judges, with a transfer-intent classifier gating hangups on transfer — so Digital Humans stop hanging up prematurely on ambiguous turns. Chat simulations gained the same discipline: conversations no longer end while the agent’s question is still pending.
NewBilling

Pay-as-you-go, auto-reload & plan limits

New orgs are provisioned straight onto pay-as-you-go with billing attached at signup, with an optional automatic credit top-up and a wallet-balance view. The billing tab was redesigned with new plan cards, credit codes are redeemable, plan switches reuse your saved card, and unused base allotment rolls over as credit. Plan entitlements — phone number counts, simulation concurrency, data retention — are now enforced in real time.
NewPlatform

Unified knowledge bases (RAG)

Knowledge bases became a full retrieval feature: upload documents, watch ingestion status, and attach knowledge bases to agents. Knowledge you upload once now grounds digital-human behavior, chat simulations, and metric judging consistently everywhere.
ImprovementPlatform

Phone number management

Buy phone numbers in bulk in a single purchase, with limits driven by your plan instead of a hardcoded cap, and a per-number SMS Enabled indicator so you know which numbers can run SMS tests.
NewObservability

Uptime Monitoring

Continuous health checks for production agents: Bluejay periodically places real calls (or SMS/webchat sessions) against your agent, detects incidents, and alerts on failures — Slack with optional @channel, a configurable consecutive-failure threshold, and latency tracking with 7-day aggregates. Health checks can run full conversations with your own Digital Human, grade with your custom metrics, respect day/hour active windows, and be paused and resumed.Uptime monitoring
NewPlatform

Voice Agent Audit

A free public audit: enter your agent’s phone number on the landing page and Bluejay runs ~20 tailored test calls with custom metrics and audio-quality checks, then emails you an exec-readable PDF report of the findings.
NewIntegration

Custom telephony connections

For orgs with their own telephony: Bluejay can call a webhook you host at dial time to fetch a per-call SIP endpoint (dynamic routing, ephemeral endpoints), LiveKit agents can authenticate via API key or a token webhook, and outbound simulations can admit authenticated inbound SIP — so setups where your infrastructure originates or routes the call are fully testable.Dynamic SIP webhook
ImprovementSimulations

No-answer diagnostics

Calls your agent never answers are now diagnosable: no-answer results carry the recording and the caller-side transcript so you can hear exactly what happened, and dial timeouts are classified correctly instead of as generic failures.
NewMetrics

Tool Call metric & tool adherence

A new tool_call metric type deterministically checks whether the agent called the right tools — pick tool names from a combobox, no LLM prompt needed. Tool adherence supports expected tool calls with parameter comparison and an order-enforcement toggle, and grading waits for the expected tools’ traces to arrive before scoring.Tool calls
NewIntegration

Vapi squads & Retell conversation flows

Multi-assistant Vapi squads render as editable workflow canvases, Retell conversation flows get a dedicated editor, and auto-sync keeps Retell, Vapi, and ElevenLabs agents current — with an org-level toggle controlling re-sync before every simulation.Vapi simulations · Retell integration
NewPlatform

Self-improve for Vapi agents

Bluejay can close the test-fix-retest loop for Vapi agents: it analyzes failed results, proposes prompt improvements, applies them, and announces runs when they finish.Self-improve
NewAlerts

Usage alerts

Get notified when your org crosses credit-usage thresholds, framed by credits remaining. Every org is seeded with default alerts at 80% and 100% of its allotment, with recipient selection including your whole team.Alerts overview
NewDigital Humans

Voice scenes

Describe a voice as a scene — “tired commuter on a train” — and Bluejay generates a matching custom voice with bundled background ambience, giving Digital Humans situationally realistic audio.Digital Humans overview
NewSimulations

Cross-agent run comparison

Compare simulation runs across different agents: stage runs into a comparison tray that persists as you browse, then open a side-by-side comparison — with PDF export and instant email send.Simulation runs
ImprovementSimulations

Simulation run controls

Finer control over how simulations run: a delay-between-tests setting, concurrency capped to your phone-number pool and org ceiling, an org-level max call length with per-simulation override, a per-simulation inactivity timeout, and scheduling restricted to chosen days and hours.
ImprovementIntegration

CHIRP playback acknowledgments

The CHIRP websocket protocol gained a mark message: agents integrating over CHIRP receive playback acknowledgments confirming when queued audio has actually been played to the caller, enabling precise turn-taking and barge-in handling.Websocket connections
NewIntegration

Google Dialogflow CX integration

Run simulations against Dialogflow CX agents end-to-end. Connect your Google Cloud project via Service Account key, point Bluejay at your CX agent, and Bluejay handles session orchestration, audio transport, and evaluation. Outbound dialing, per-agent CES region, and CX-specific dispatch are all supported.Dialogflow CX simulations
NewIntegration

Kore.ai integration

Bluejay now integrates with Kore.ai for both simulations and observability. Connect your Kore platform host, bot ID, and webhook app credentials on the agent’s connection settings to start testing Kore-powered conversational agents.Kore.ai integration
NewIntegration

Retell outbound support

Retell-connected agents can now run as outbound. Bluejay auto-derives the agent’s type (inbound vs. outbound) from your Retell bindings on sync, provisions always-on SIP for outbound Digital Humans, and preserves the synced phone number across connection saves. Outbound Digital Humans default to not speaking first.Retell integration
ImprovementObservability

Retell observability + simulations refresh

We rewrote the Retell Observability page around the new UI flow and expanded the Retell Simulations guide with quick-sync steps, metrics coverage, and updated walkthroughs.Retell Observability · Retell Simulations
NewIntegration

Vapi simulations

Run simulations directly against Vapi assistants. Bluejay connects to your Vapi assistant, drives the conversation with a Digital Human, and evaluates the resulting transcript with your custom metrics.Vapi simulations
NewObservability

Audio quality metrics (Tyto risk score)

Call logs now include a Tyto risk_score as a headline audio-quality signal alongside latency and custom metrics. The score is plumbed end-to-end through the eval DTOs and aggregated on the call logs view so you can sort and filter production conversations by audio risk.Observability overview
NewAlerts

PagerDuty alert channel

Threshold alarms can now page on-call via PagerDuty Events API v2 in addition to Slack, email, and webhook. Connect your PagerDuty integration key on the notifications page and route specific alarms to PagerDuty for incident-grade escalation.Alerts overview
ImprovementAlerts

Percentage alarms + simulation alerts

Threshold alarms now support percentage-based thresholds (e.g. alert when goal success drops below 80% of calls) in addition to absolute values, and you can scope alarms to simulation metrics as well as observability. Percentage alarms gracefully recover when no agents have calls in the window.Alerts overview
NewSimulations

Customer Journeys

Customer Journeys let you chain multiple simulated calls into a single multi-step scenario, with prior-call context (entities, knowledge, and state) carried across steps. Build journeys from the dashboard or define them in Bluejay-as-Code.Customer Journeys guide · Bluejay-as-Code
NewDigital Humans

Voice cloning in simulations

Digital Humans can now speak with a cloned voice during simulations. Upload a voice sample on the Digital Human, and Bluejay will use that cloned voice for all simulation runs — useful for accent coverage, brand-voice testing, and replicating known caller profiles.Digital Humans overview
NewSimulations

Custom background noise

Upload your own background noise audio to play under simulation calls. Test how your agent handles call-center hum, traffic, household noise, or any acoustic environment you care about, instead of relying only on the built-in profiles.Simulations overview
NewDigital Humans

Bulk CSV upload for Digital Humans

Create Digital Humans in bulk by uploading a CSV from the Digital Humans page. Each row maps to one Digital Human (persona, scenario, traits, and success criteria), so you can stand up large evaluation populations without scripting against the API.Digital Humans overview
NewVoices

New voice accents

Expanded voice coverage for Digital Humans:
  • UK English — Scottish, Irish, and Welsh accents
  • Korean — new locale support
Voices resolve from atomic language, accent, and gender attributes, so existing agents pick up new options automatically.
ImprovementDocs

Scenario Builder split from Workflows

We split Scenario Builder (single-agent conversational scripts) from Workflows (multi-agent automation pipelines) across the docs and API reference. The Workflows API was renamed to the Scenario Builder API where it referred to single-agent scenarios; multi-agent workflow endpoints stay under Workflows.Scenario Builder overview · Workflows overview
NewReference

Simulation Statuses reference

Added a dedicated Simulation Statuses reference page covering every terminal status, the error codes that map to each failure bucket, and what to do when a simulation lands in System Error. Useful when triaging failed runs or building dashboards on top of the API.Simulation statuses
NewDigital Humans

DTMF tool for Digital Humans

Digital Humans can now be configured to send DTMF tones during a call via a new allow_dtmf_tool field, aligned with how allow_end_call_tool and allow_silence_tool are stored and returned. Create and bulk-create default to allow_dtmf_tool: true when omitted; updates treat the field as optional (omit to leave unchanged). OpenAPI schemas DigitalHumanRequestData, DigitalHumanResponseData, and UpdateDigitalHumanRequest include the new property.Create Digital Human · Update Digital Human
NewSimulations

Runs per Digital Human

Simulations now support executing multiple runs for each selected Digital Human in a single batch. Set runs_per_digital_human on the simulation (persisted in experiments.settings) to control the default, or pass an explicit per-run override when queuing a run. Values must be positive integers; when unset, each Digital Human runs once. The Bluejay dashboard exposes the new control on the simulation settings page and the “Create simulation” and “Create new run” dialogs.Simulations overview
ImprovementDigital Humans

Digital Human intelligence upgrade

We upgraded the reasoning quality of Digital Humans during both simulations and observability replays. Expect more coherent multi-turn behavior, better adherence to persona and objectives, and steadier tool-use decisions across longer conversations. No configuration changes are required — the upgrade applies automatically to all Digital Humans.Digital Humans overview
ImprovementWorkflows

Scripted silence in workflows

When a Digital Human is running in workflow mode, the silence tool now evaluates only at end-of-turn instead of firing from timeout-based checks mid-turn. This removes a class of false silence triggers on scripted workflow steps and keeps silence decisions aligned with the workflow’s turn boundaries. Non-workflow Digital Humans are unchanged.Workflows overview
PerformanceWorkflows

Workflow latency fix

Fixed a latency regression in workflow-mode conversations where session-handler state could delay turn transitions. Workflow runs now advance between agent and user turns without the extra wait, improving perceived responsiveness on branching paths.Workflows overview
ImprovementAPI

Digital Human: silence tool fields

Digital Human create, read, update, delete, and bulk APIs now expose allow_silence_tool and silence_tool_instructions, aligned with how allow_end_call_tool and hangup_instructions are stored and returned. Create and bulk-create default to allow_silence_tool: false and silence_tool_instructions: "default" when omitted; updates treat fields as optional (omit to leave unchanged). OpenAPI schemas DigitalHumanRequestData, DigitalHumanResponseData, and UpdateDigitalHumanRequest include the new properties.Create Digital Human · Update Digital Human
ImprovementAPI

Redesigned Workflows

Workflows are a structured way to test your agent along a conversation path.
  • Agent and user turns — you define what the agent should say or do (so simulations can check it) and what the Digital Human says on each step.
  • Branching — options nodes capture different things the caller might do next (speech, DTMF, silence).
  • Coverage — each distinct path through the graph becomes its own Digital Human, so every branch gets exercised.
  • Docs — cookbook and API reference refreshed; older workflow endpoints are under Deprecated.
Create scenario · Cookbook
ImprovementDocs

Revamped documentation

We completely overhauled the Bluejay docs with improved navigation, expanded guides, and new content across the board. Highlights include:
  • Restructured navigation with dedicated tabs for Documentation, API Reference, and Changelog
  • Expanded integration guides covering all supported providers and simulation transports
  • Cookbook recipes for common workflows like GitHub Actions CI, API-driven evaluations, and webhook setup
Explore the docs
NewSimulations

SMS simulation support

Bluejay now supports SMS-based simulations, enabling you to test text-based agent flows end-to-end. Configure SMS simulations the same way you configure voice simulations — define Digital Humans, set up custom metrics, and run batch evaluations against your SMS agent.SMS simulations support all existing integrations including telephony providers and HTTP webhooks.Read the docs
NewAlerts

Threshold alarms

Set threshold-based alarms on any custom metric to get alerted when agent performance degrades. Define upper or lower bounds, choose your notification channel (Slack, email, or webhook), and Bluejay will trigger alerts automatically when production metrics cross your thresholds.Alarms work across both observability and simulation metrics, so you can catch regressions in production and in testing.Read the docs
PerformanceObservability

Faster call log evaluation

We made significant performance improvements to the observability evaluation pipeline:
  • 3x faster evaluation for call logs with custom metrics
  • Parallel metric execution — multiple custom metrics now evaluate concurrently instead of sequentially
  • Reduced API latency for the /evaluate and /re-evaluate endpoints by approximately 40%
These improvements apply automatically to all existing observability configurations.Read the docs
NewIntegration

Miro integration for simulation visualization

You can now connect Bluejay to Miro to automatically generate visual conversation flow diagrams from your simulation results. Each simulation run produces a Miro board showing the conversation tree, branching paths, and metric outcomes.Connect your Miro workspace from the Integrations page in your Bluejay dashboard.Read the docs
ImprovementWorkflows

Workflow scheduling improvements

Workflows now support cron-based scheduling with finer granularity. You can schedule simulation runs and observability evaluations to execute at specific intervals — hourly, daily, or on a custom cron expression.Additional improvements include:
  • Retry logic for failed workflow steps
  • Execution history with detailed step-level logs
  • Webhook notifications on workflow completion or failure
Read the docs
NewSimulations

Community-based simulation runs

You can now run simulations against an entire community of Digital Humans in a single batch. Previously, simulations ran against individual Digital Humans or manually selected groups. Community-based runs let you test your agent against a diverse, pre-configured population in one click.Combine communities with custom metrics to get aggregate performance scores across demographic segments, persona types, or behavioral profiles.Read the docs
ImprovementMetrics

Custom metric formulas

Custom metrics now support formula-based definitions in addition to LLM-as-a-Judges. Define metrics using arithmetic expressions over existing metric scores, enabling composite metrics like weighted averages or pass/fail thresholds without writing evaluation prompts.Formula metrics evaluate instantly and do not consume LLM credits.Read the docs
NewIntegration

Pipecat integration

Bluejay now integrates natively with Pipecat for running simulations against Pipecat-powered voice agents. Connect your Pipecat pipeline endpoint and Bluejay will handle session orchestration, audio transport, and evaluation.Read the docs
NewAPI

Knowledge base versioning API

The Knowledge Base API lets you create knowledge bases, add documents, and attach them to agents. Knowledge bases ground your agents’ answers and power hallucination detection during evaluation.Read the docs
ImprovementDashboard

Dashboard redesign

We redesigned the Bluejay dashboard with a focus on surfacing actionable insights. The new layout includes:
  • At-a-glance health scores for each agent across simulation and production metrics
  • Trend sparklines showing metric performance over time
  • Alert badges highlighting agents that need attention
  • Quick-launch actions for running simulations and viewing recent call logs
The redesign is live for all users.Read the docs
NewIntegration

ElevenLabs observability integration

Bluejay now supports direct observability integration with ElevenLabs Conversational AI. Connect your ElevenLabs account to automatically ingest call logs, evaluate them with custom metrics, and surface quality issues in your dashboard.Read the docs
ImprovementSimulations

Digital Human generation improvements

The Digital Human generation engine has been upgraded with better persona diversity and more realistic conversational styles:
  • Expanded trait library with 40+ new customer traits including emotional tone, technical proficiency, and communication preferences
  • Scenario-aware generation — Digital Humans now adapt their behavior based on the simulation scenario context
  • Bulk generation — generate up to 100 Digital Humans in a single API call
Read the docs
NewIntegration

Slack alerting integration

Connect Bluejay to Slack to receive real-time alerts when production metrics drop below thresholds or simulation runs complete. Configure per-channel routing so the right team gets the right alerts.Read the docs
NewSimulations

WebSocket simulation support

Bluejay now supports WebSocket-based simulations for testing real-time, bidirectional agent communication. Configure your WebSocket endpoint, define the message protocol, and run simulations with full transcript capture and metric evaluation.Read the docs
ImprovementAPI

Prompt versioning and labels

The Prompt API now supports versioning and labeling. Create multiple versions of a prompt, tag them with labels like production or staging, and reference them by label in your agent configuration. Roll back to any previous version instantly.Read the docs
NewObservability

Webhook-based log ingestion

You can now send call logs to Bluejay via webhook for evaluation. Configure a webhook endpoint in your dashboard, point your agent platform at it, and Bluejay will automatically ingest, evaluate, and store the results.This is the fastest way to get observability running if your platform isn’t covered by a native integration.Read the docs
NewIntegration

LiveKit simulation integration

Run simulations against LiveKit-powered voice agents. Bluejay connects to your LiveKit room, manages participant sessions, and captures full audio transcripts for evaluation.Read the docs
NewMetrics

Metrics Lab

Introducing Metrics Lab — an interactive environment for prototyping and testing custom metrics before deploying them. Write evaluation prompts, test them against sample transcripts, and iterate on scoring criteria without affecting production data.Read the docs
ImprovementDashboard

Folder-based agent organization

Agents can now be organized into folders for better workspace management. Create folders, move agents between them, and filter your agent list by folder. Folders are available in both the dashboard and the API.Read the docs
NewIntegration

Retell observability integration

Bluejay now integrates directly with Retell for production call monitoring. Connect your Retell account to automatically pull call logs, run evaluations, and track agent performance over time.Read the docs
NewIntegration

Vapi observability integration

Bluejay now integrates with Vapi for production call monitoring. Connect your Vapi account to automatically ingest call logs, run custom metric evaluations, and track agent quality over time.Read the docs
NewIntegration

Bland observability integration

You can now connect Bluejay to Bland for production observability. Call logs from your Bland-powered agents are automatically ingested, evaluated against your custom metrics, and surfaced in the dashboard.Read the docs
NewAPI

Communities and workflows API endpoints

New API endpoint groups for managing communities and workflows:
  • Communities — create, update, add members, list, and delete communities programmatically
  • Workflows — define, schedule, and manage automation workflows through the API
Read the docs
NewIntegration

SIP simulation integration

Bluejay now supports SIP-based simulations. Connect your SIP trunk and run simulations directly over the SIP protocol, enabling testing for enterprise telephony deployments and contact center agents.Read the docs
NewIntegration

Telephony simulation support

Run simulations over PSTN by connecting your telephony provider to Bluejay. Dial into your agent’s phone number, capture the full conversation, and evaluate it with custom metrics — all without changing your agent’s infrastructure.Read the docs
NewAPI

Observability and evaluation API endpoints

New API endpoint groups for observability and evaluation workflows:
  • Observability — evaluate and re-evaluate call logs, manage call log lifecycle
  • Custom Metrics — create, bulk-create, update, list, and delete custom metrics via API
Read the docs
NewAPI

Digital Humans and simulation runs API endpoints

New API endpoint groups for simulation orchestration:
  • Digital Humans — create, generate, update, list, and manage Digital Human personas
  • Simulation Runs — queue voice and SMS runs, retrieve results, and manage active conversations
Read the docs
NewAPI

Agents and simulations API endpoints

The first set of public API endpoints is now available:
  • Agents — create, update, list, move, and delete agents
  • Simulations — create, configure, list, and manage simulations programmatically
These endpoints form the foundation of the Bluejay API and enable full automation of your testing pipeline.Read the docs
NewSimulations

Digital Human personas

Introducing Digital Humans — synthetic customer personas that power Bluejay simulations. Define demographic profiles, personality traits, communication styles, and scenario-specific behaviors to create realistic test conversations at scale.Read the docs
NewSimulations

Simulation engine

The Bluejay simulation engine is live. Run synthetic conversations against your voice agents to validate behavior before production. Define scenarios, assign Digital Humans, and evaluate performance with custom metrics.Read the docs
NewObservability

Observability pipeline

Bluejay’s observability pipeline is now available. Ingest production call logs, evaluate them against custom metrics, and surface quality trends in the dashboard. Supports both API-based and webhook-based log ingestion.Read the docs
NewMetrics

Custom metrics engine

Define custom evaluation criteria tailored to your use case. Bluejay’s custom metrics engine supports LLM-as-a-Judge evaluations with configurable scoring rubrics, pass/fail thresholds, and dynamic variables that adapt to conversation context.Read the docs
NewDashboard

Agent management and dashboard

The Bluejay dashboard is live. Create and manage your conversational AI agents, view performance summaries, and navigate your workspace from a central hub.Read the docs
NewPlatform

Evaluation framework

Bluejay’s evaluation framework is ready. Score agent conversations using structured rubrics, capture per-turn and per-call metrics, and generate evaluation reports. This framework underpins both simulation testing and production observability.
NewPlatform

Core platform infrastructure

The foundational Bluejay platform is up and running — authentication, workspace management, and the base API layer. This milestone sets the stage for all product features to follow.