Skip to content

[Epic] Analytics, conversation intelligence, evaluations, and experiments #80

Description

@rahuliitk

Parent roadmap: #76

Problem

QuickVoice has useful KPI, volume, performance, call-detail, transcript, extracted-data, and evaluation surfaces, while several open PRs propose Langfuse tracing. It does not yet provide one stable metric model, end-to-end operational and business reporting, searchable conversation intelligence, systematic regression evaluation, or version-aware production experimentation.

Outcome

Operators can understand what happened and why; builders can convert failures into repeatable tests; reviewers can combine automated and human quality signals; and owners can promote a better agent version using statistically and operationally defensible evidence.

Required capabilities

  • Canonical organization-scoped events for calls, sessions, workflow nodes, turns, model/provider spans, tools, transfers, campaign actions, evaluations, and business outcomes.
  • Clear metric definitions and freshness, timezone, currency, filter, and aggregation semantics.
  • Dashboards for outcome, funnel, cohort, cost, latency, reliability, workflow paths, campaigns, tools, providers, models, and agents.
  • Keyword and semantic transcript search, topics, summaries, sentiment, intent, objections, customer timeline, and saved views.
  • Scenario datasets, tests generated from calls, next-response/tool/full-conversation assertions, controlled mocks, and simulated users.
  • Repeated probabilistic runs, failure clustering, flakiness, baseline comparison, human QA queues, reviewer calibration, and release gates.
  • Immutable configuration versions, branches, deterministic traffic allocation, canaries, promotion, rollback, and experiment analysis.

Child issues

Dependencies

  • Workflow/version contracts.
  • Customer/campaign identity model.
  • Public API/event and privacy contracts.
  • Consolidated observability implementation from the overlapping Langfuse PRs.

Acceptance criteria

  • A documented metric produces the same result in dashboard, API, export, and experiment analysis for the same filters.
  • Every trace and stored artifact is organization scoped and applies zero-PII/retention policy before export to optional observability providers.
  • Users can drill from an aggregate anomaly to relevant calls, spans, workflow nodes, transcripts, and evaluation evidence where policy permits.
  • A production failure can become a sanitized regression test without copying protected data by default.
  • Release gates can block promotion but require an authorized, reasoned, audited override path.
  • Experiment assignment is stable, mutually exclusive, observable, and reversible.
  • Metric correctness, authorization, timezone/currency, redaction, evaluation determinism boundaries, statistical summaries, and rollback are tested.

Boundary

Automated sentiment, summaries, and LLM judgments are assistive signals with confidence and provenance, not unquestionable facts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: ai-runtimePython AI API and LiveKit workerarea: consoleCustomer consolearea: serverExpress API and server control planeenhancementNew feature or requeststatus: needs-designNeeds maintainer design agreement before implementation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions