AI Agents
Amplitude Agent Analytics Ties AI Agent Performance Directly to Business Outcomes
The platform provides full-session scoring on task completion, quality and safety while linking agent data to product events, enabling enterprises to quantify effects on retention and revenue as demonstrated in production deployments.
Amplitude Agent Analytics is a production monitoring platform that bridges LLM observability and product analytics to tie AI agent session quality directly to business outcomes like retention and revenue.
Enterprises deploying AI agents in production have long faced the challenge of moving beyond sampled manual reviews to comprehensive measurement of every interaction. Amplitude Agent Analytics addresses this gap by instrumenting full session traces as queryable events that sit alongside existing product data within the same analytics environment. This integration allows teams to examine whether agent performance improvements translate into measurable lifts in user retention or revenue without requiring separate data pipelines.
Prior approaches relied on limited sample sets that often missed systemic failure modes across thousands of daily sessions. With Agent Analytics, every closed session receives automatic scoring on multiple built-in signals, each accompanied by a written rationale generated through either default evaluators or custom LLM-as-a-judge and code-based rules. The system organizes records into sessions, individual turns and agent instances, then clusters user intents and failure modes to surface patterns that engineers can address systematically.
What background led to the need for dedicated AI agent analytics platforms?
Production AI agents operate across multi-turn conversations that involve retrieval, reasoning and response generation, creating complex traces that traditional observability tools struggle to contextualize against business results. Design partners including The Economist tested early versions of Agent Analytics against live agents and identified the requirement for automatic instrumentation that attaches signals at ingestion time rather than through post-hoc analysis. This shift eliminated the previous reliance on drawing conclusions from handfuls of sessions and enabled continuous monitoring at scale.
The availability of the tool to all customers, including those on the free plan, occurred in June 2026 after iterative testing that refined clustering algorithms and evaluator accuracy. Enterprises ranging from early-stage startups to established media organizations validated performance across diverse agent architectures, confirming that the same data model supports both cost optimization experiments and quality improvement initiatives without additional configuration overhead.
How does Agent Analytics score sessions and cluster failure modes in practice?
Upon session closure the platform applies a standardized set of evaluators covering task completion, response quality, alignment with user intent, overall session safety, presence of user friction, receipt of negative feedback and data quality issues. Each evaluation produces both a quantitative score and a natural-language rationale that explains the assessment, allowing teams to understand not only what occurred but why the system assigned a particular rating. Custom evaluators can supplement the built-in signals when domain-specific criteria are required.
Clustered failure modes surface recurring issues such as repeated retrieval errors or intent misclassifications across large volumes of sessions. These clusters feed directly into ticket creation workflows, routing specific failure categories to the appropriate engineering or product teams. Because the underlying data remains queryable as standard Amplitude events, analysts can apply the same cohort and funnel analyses used for product metrics to agent performance dimensions.
| Metric | Before Implementation | After Implementation | Reported Change |
|---|---|---|---|
| Task success rate | Sampled reviews only | 96.9 percent | Full population measurement enabled |
| Weekly task failures | Baseline volume | Reduced 84 percent | Systematic issue resolution |
| Session cost per active user | 4.88 dollars | 2.33 dollars | 52.3 percent reduction |
| Conversion rate | Approximately 30 percent | Approximately 30 percent | Flat while cost declined |
What technical specifics enable integration between agent traces and product data?
Agent sessions inherit the same user identifier and project context used for product events, eliminating the need for custom joins or external data movement. This shared schema means that a successful agent interaction can be directly correlated with subsequent conversion events, renewal likelihood or revenue figures within the same analysis environment. Auto-instrumentation ensures that trace details such as retrieved documents, model responses and evaluator scores appear as structured properties on each event record.
Support for both LLM-as-a-judge and code-based evaluators allows teams to define quality thresholds that align with internal standards. When thresholds are breached the system can trigger automated alerts or ticket filing. The resulting dataset supports longitudinal analysis, revealing whether improvements in task success rates correspond to higher user retention or increased lifetime value over multiple quarters.
- Instrument every agent session with built-in and custom evaluators at ingestion time
- Cluster intents and failure modes to prioritize engineering effort
- Auto-file tickets for recurring issues identified through clustering
- Share user and project identifiers with product event streams for joint analysis
- Query agent traces using the same cohort and funnel tools applied to product metrics
- Track cost, latency and quality changes across model versions in production
How did The Economist apply Agent Analytics to its Lens research assistant?
The Economist Group's Lens AI research assistant on the Viewpoint platform transitioned from sampled manual reviews to complete session measurement after adopting Agent Analytics. The multi-turn agent now logs every interaction with attached signals, enabling the team to identify specific failure categories and measure the impact of targeted fixes. This full-signal approach replaced earlier reliance on limited session samples that had constrained visibility into overall performance.
By Amplitude's measure, Lens now holds a 96.9 percent task success rate, and weekly task failures fell 84 percent as we worked through the issues the data surfaced. Those gains came from ordinary engineering. The measurement is what tells us where to point it.Fionn O'Raghallaigh, AI Group Product Manager, The Economist
Fionn O'Raghallaigh noted that prior to Agent Analytics the team lacked granular visibility into individual turns, retrieved content and scoring details across the full session population. The new instrumentation allowed direct inspection of each turn within a session, revealing exactly where retrieval or response generation introduced errors. This level of detail accelerated root-cause analysis and supported data-driven decisions on model selection and prompt refinement.
What are the market and stakeholder implications for enterprises scaling AI agents?
Organizations that link agent performance to product metrics gain the ability to justify continued investment in AI infrastructure through demonstrated effects on retention and revenue. The shared data model reduces the friction between AI engineering teams and product analytics groups, fostering unified roadmaps that treat agent quality as a core product metric rather than an isolated technical concern. This alignment is particularly relevant for enterprises managing multiple agent deployments across customer support, research and internal productivity use cases.
The availability of the platform on free plans lowers the barrier for smaller teams to establish baseline measurement before scaling. Larger enterprises benefit from the same tooling used by design partners, ensuring consistency in evaluation standards across the ecosystem. As agent usage grows, the ability to quantify the business impact of each incremental improvement in task success or failure reduction becomes a competitive differentiator in resource allocation decisions.
What expert reactions emphasize the value of full-signal measurement?
Fionn O'Raghallaigh highlighted the transition from drawing conclusions from limited sessions to accessing comprehensive data that surfaces actionable patterns. The second quotation from the same source underscores the practical benefit of viewing each turn, retrieved content and assigned score within a single interface. These observations illustrate how complete instrumentation changes the operational tempo from reactive debugging to proactive optimization.
Industry observers note that the shift from sampled to exhaustive measurement mirrors earlier evolutions in web and mobile analytics, where full event capture enabled more precise attribution of user behavior to business outcomes. Agent Analytics extends this principle to conversational AI, providing the measurement foundation required for sustained production deployments.
What developments are expected next for AI agent analytics capabilities?
Continued expansion of custom evaluator options and tighter integration with existing ticketing and monitoring systems are anticipated as adoption widens. The ability to run comparative analyses across model versions while holding user cohorts constant will support more rigorous cost-quality tradeoff decisions. Enterprises are also expected to extend the same session-level scoring frameworks to multi-agent orchestration scenarios where several specialized agents collaborate within a single user journey.
As regulatory and compliance requirements evolve, the structured rationales produced by evaluators may serve as audit artifacts demonstrating that agent behavior meets defined safety and quality thresholds. The combination of quantitative scores and natural-language explanations provides both the metrics and the context needed for internal governance and external reporting.
Further enhancements in clustering granularity and automated ticket routing will reduce the manual effort required to translate measurement insights into engineering action. These capabilities collectively position agent analytics as a standard component of production AI stacks rather than an optional add-on.
Frequently asked
How does Amplitude Agent Analytics differ from traditional LLM observability tools?
It automatically scores every session on task completion, quality and safety while sharing user identifiers with product events, enabling direct correlation of agent performance with business metrics such as retention and revenue.
What results did The Economist achieve after adopting the platform?
The Lens assistant reached 96.9 percent task success and reduced weekly failures by 84 percent through systematic issue resolution guided by full-session data.
When did Agent Analytics become generally available?
The tool became available to all customers, including free-plan users, in June 2026 following design-partner testing with enterprises such as The Economist.
Sources
- Amplitude — Agent Analytics is now available to all customers including the free plan after design partner testing; session cost per active user dropped from 4.88 dollars to 2.33 dollars representing a 52.3 percent reduction while conversion remained flat near 30 percent.
- Amplitude — The Economist Group used Amplitude Agent Analytics on its Lens AI assistant, reaching 96.9 percent task success and cutting weekly task failures by 84 percent; sessions landed into the platform with signals attached out of the box.
- Amplitude — Agent Analytics has been tested by customers ranging from Y Combinator startups to enterprise companies such as The Economist and supports full production monitoring of multi-turn agents.