End-to-end traces
Follow model calls, tools, errors, latency, and usage through the complete run that produced them.
A trace explains what happened. An evaluation tells you whether it met the expectation. Lens keeps both connected so production evidence can inform the next release.
Follow model calls, tools, errors, latency, and usage through the complete run that produced them.
Keep suite, case, metric, release, dataset version, and trace identity together as one evaluation run.
Compare a candidate with its baseline and make regressions visible before they reach production.
Send native Anvia telemetry or supported OTLP traces to infrastructure and storage your team operates.
$docker compose up -dStart with what happened in production, diagnose the behavior, measure the candidate on stable cases, then let an explicit gate carry the evidence into release review.
Capture the agent, model, tool, release, latency, and usage involved in one production run.
Move from an alert or regression into the exact span, input boundary, and failure that changed.
Replay stable cases against a published dataset version and correlate every result to its trace.
Compare the candidate with its baseline, apply thresholds, and ship only with evidence.
Explore an example overview with sample data. Change the time window and inspect token usage, latency, model efficiency, and tool health.
Production AgentsOverview
448 spans · 4 active users in this window
Input and output tokens over time
Traces and LLM generations, with failed traces highlighted
P50 and P95 duration for generation observations
Share of total generation tokens
Usage, duration, and reliability by generation model
| Model | Generations | Token share | Input | Output | Total | Tokens / gen | P95 duration | Errors |
|---|---|---|---|---|---|---|---|---|
| gemini-2.5-flash | 48 | 25.5% | 29.3K | 23K | 52.3K | 1,089 | 1.76s | 2.1% |
| gpt-4.1 | 48 | 25.5% | 29.3K | 23K | 52.3K | 1,089 | 1.86s | 4.2% |
| gpt-4.1-mini | 48 | 25.4% | 29.2K | 22.9K | 52.1K | 1,085 | 1.58s | 4.2% |
| claude-sonnet-4 | 48 | 23.6% | 27.1K | 21.3K | 48.4K | 1,008 | 1.92s | 6.3% |
Token load and trace health by service
| Service | Traces | Generations | Tokens | P95 duration | Errors |
|---|---|---|---|---|---|
| customer-support-agent | 28 | 86 | 94.3K | 3.12s | 9.4% |
| billing-copilot | 14 | 40 | 59.4K | 1.96s | 6.1% |
| research-agent | 11 | 33 | 30.8K | 4.80s | 14.2% |
| docs-qa | 11 | 33 | 20.5K | 1.44s | 13.0% |
Most-used tool calls, duration, and failures
| Tool | Calls | P95 call duration | Errors |
|---|---|---|---|
| search_knowledge_base | 96 | 428ms | 2.1% |
| get_customer_context | 64 | 312ms | 0.0% |
| lookup_invoice | 32 | 186ms | 3.1% |
Highest token usage in this window
Latest traces with an error
Lens groups the suite, immutable dataset version, metric direction, usage, outcomes, and trace references as one candidate that can be compared and gated.
Dataset support-regression@v12 · 48 cases · 192 metric results
@anvia/lens creates isolated telemetry providers, attaches to the runtime you already own, and exports correlated traces over OTLP HTTP with project-scoped credentials.
1import { Agent } from '@anvia/core'2import { LensClient } from '@anvia/lens'34const lens = new LensClient()5const tracing = lens.observer({ captureMode: 'safe' })67const agent = new Agent({8 id: 'support',9 model,10 observability: { observers: { tracing } },11})1213await agent.generate({ prompt: 'Summarize this ticket.' })14await lens.flush()15await lens.close()Safe capture exports operational metadata without input or output bodies. Full payloads remain opt-in; redaction and size limits are configurable and governed by your application.
Defines the agent, release policy, credentials, and which environments or data classes may include payloads.
No global provider takeover. Payload capture is off until explicitly enabled.
Agent, model, tool, environment, release, latency, usage, and errors.
Run lifecycle, cases, metrics, outcomes, usage, and direct trace references.
Store project-scoped traces and evaluation evidence in your Lens deployment.
Use flush() for an explicit delivery checkpoint; call close() during final cleanup.
Three repeatable paths from an operational or quality signal to the run-level evidence needed to act.
Compare P95 by release and service, open the slow cohort, then inspect the model and tool spans consuming the budget.
Group errors by service and tool, follow one representative run, and separate bad input from permissions, retries, or provider failure.
Open the candidate comparison, inspect the failing cases, jump to their traces, then rerun the same published dataset version.
Studio helps you understand an agent before it ships. Lens is an optional production companion that keeps the same runtime concepts visible after deployment. Your agent can also use another observability backend.
Exercise the real local agent, resolve approvals, inspect context, and replay the workflow.
Explore Studio ProductionInvestigate deployed runs, compare stable evaluation suites, and carry evidence into release review.
Open the Lens guideSame agent and trace concepts, from your machine to production.
Download the Compose and environment templates, configure unique secrets, then start the stack. The first account becomes the workspace owner.
Download the official templates, pin a Lens release, and configure unique secrets.
Launch the stack. The first account created becomes the workspace owner.
Create a project and ingestion key, then send native Anvia telemetry or supported Langfuse v5 OTLP traces.
Connect one agent in safe-capture mode, investigate its traces, then promote a stable evaluation suite into an explicit quality gate.