You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
QuickVoice has useful KPI, volume, performance, call-detail, transcript, extracted-data, and evaluation surfaces, while several open PRs propose Langfuse tracing. It does not yet provide one stable metric model, end-to-end operational and business reporting, searchable conversation intelligence, systematic regression evaluation, or version-aware production experimentation.
Outcome
Operators can understand what happened and why; builders can convert failures into repeatable tests; reviewers can combine automated and human quality signals; and owners can promote a better agent version using statistically and operationally defensible evidence.
Required capabilities
Canonical organization-scoped events for calls, sessions, workflow nodes, turns, model/provider spans, tools, transfers, campaign actions, evaluations, and business outcomes.
Clear metric definitions and freshness, timezone, currency, filter, and aggregation semantics.
Dashboards for outcome, funnel, cohort, cost, latency, reliability, workflow paths, campaigns, tools, providers, models, and agents.
Keyword and semantic transcript search, topics, summaries, sentiment, intent, objections, customer timeline, and saved views.
Scenario datasets, tests generated from calls, next-response/tool/full-conversation assertions, controlled mocks, and simulated users.
Repeated probabilistic runs, failure clustering, flakiness, baseline comparison, human QA queues, reviewer calibration, and release gates.
Parent roadmap: #76
Problem
QuickVoice has useful KPI, volume, performance, call-detail, transcript, extracted-data, and evaluation surfaces, while several open PRs propose Langfuse tracing. It does not yet provide one stable metric model, end-to-end operational and business reporting, searchable conversation intelligence, systematic regression evaluation, or version-aware production experimentation.
Outcome
Operators can understand what happened and why; builders can convert failures into repeatable tests; reviewers can combine automated and human quality signals; and owners can promote a better agent version using statistically and operationally defensible evidence.
Required capabilities
Child issues
Dependencies
Acceptance criteria
Boundary
Automated sentiment, summaries, and LLM judgments are assistive signals with confidence and provenance, not unquestionable facts.