The feedback loop for production agents.

A trace explains what happened. An evaluation tells you whether it met the expectation. Lens keeps both connected so production evidence can inform the next release.

End-to-end traces

Follow model calls, tools, errors, latency, and usage through the complete run that produced them.

Comparable evaluation

Keep suite, case, metric, release, dataset version, and trace identity together as one evaluation run.

Evidence-backed gates

Compare a candidate with its baseline and make regressions visible before they reach production.

Self-hosted control

Send native Anvia telemetry or supported OTLP traces to infrastructure and storage your team operates.

$docker compose up -d

Turn one run into the next decision.

Start with what happened in production, diagnose the behavior, measure the candidate on stable cases, then let an explicit gate carry the evidence into release review.

  1. 01

    Trace

    Capture the agent, model, tool, release, latency, and usage involved in one production run.

  2. 02

    Diagnose

    Move from an alert or regression into the exact span, input boundary, and failure that changed.

  3. 03

    Evaluate

    Replay stable cases against a published dataset version and correlate every result to its trace.

  4. 04

    Gate

    Compare the candidate with its baseline, apply thresholds, and ship only with evidence.

The operating picture, in one place.

Explore an example overview with sample data. Change the time window and inspect token usage, latency, model efficiency, and tool health.

https://lens.local/production-agents

Production AgentsOverview

Overview

448 spans · 4 active users in this window

Sample data
Total tokens205K↑ 1,342.3% vs previous period
Total cost$0.6174↑ 1,253.3% vs previous period
Tokens / generation1,066↓ 32.4% vs previous period
Active models4↑ 33.3% vs previous period
Traces64↑ 2,033.3% vs previous period
Error rate10.9%↑ 10.9 pp vs previous period
P95 generation duration1.78s↓ 10.6% vs previous period
Active sessions22↑ 1,000.0% vs previous period

Token usage

Input and output tokens over time

07 PM11 PM03 AM07 AM11 AM03 PM06 PM
  • Input tokens
  • Output tokens

Throughput and errors

Traces and LLM generations, with failed traces highlighted

07 PM11 PM03 AM07 AM11 AM03 PM06 PM
  • Generations
  • Traces
  • Errors

Generation duration

P50 and P95 duration for generation observations

07 PM11 PM03 AM07 AM11 AM03 PM06 PM
  • P50
  • P95

Tokens by model

Share of total generation tokens

015K30K45K60K

Model efficiency

Usage, duration, and reliability by generation model

ModelGenerationsToken shareInputOutputTotalTokens / genP95 durationErrors
gemini-2.5-flash4825.5%29.3K23K52.3K1,0891.76s2.1%
gpt-4.14825.5%29.3K23K52.3K1,0891.86s4.2%
gpt-4.1-mini4825.4%29.2K22.9K52.1K1,0851.58s4.2%
claude-sonnet-44823.6%27.1K21.3K48.4K1,0081.92s6.3%

Services

Token load and trace health by service

ServiceTracesGenerationsTokensP95 durationErrors
customer-support-agent288694.3K3.12s9.4%
billing-copilot144059.4K1.96s6.1%
research-agent113330.8K4.80s14.2%
docs-qa113320.5K1.44s13.0%

Tool health

Most-used tool calls, duration, and failures

ToolCallsP95 call durationErrors
search_knowledge_base96428ms2.1%
get_customer_context64312ms0.0%
lookup_invoice32186ms3.1%

Token-heavy traces

Highest token usage in this window

Resolve a customer support requestgpt-4.1 · 2m ago7,230tokens
Summarize the research findingsgemini-2.5-flash · 5m ago6,418tokens
Explain an invoice adjustmentclaude-sonnet-4 · 8m ago5,804tokens

Recent failures

Latest traces with an error

Retrieve account permissionsgpt-4.1-mini · 3m ago1,240tokens
Search the support knowledge basegemini-2.5-flash · 11m ago2,186tokens

A score describes a case. A gate informs a release.

Lens groups the suite, immutable dataset version, metric direction, usage, outcomes, and trace references as one candidate that can be compared and gated.

Blocked
Candidate

support-agent@1.4.0-candidate.2

Dataset support-regression@v12 · 48 cases · 192 metric results

Passed
46
Failed
2
Eval usage
184k
MetricBaselineCandidateDelta
Answer relevancyHigher is better0.910.94+0.03pass
FaithfulnessHigher is better0.960.90−0.06fail
HallucinationLower is better0.040.07+0.03fail
Turn relevancyHigher is better0.880.90+0.02pass

Observe the agent. Flush the evidence.

@anvia/lens creates isolated telemetry providers, attaches to the runtime you already own, and exports correlated traces over OTLP HTTP with project-scoped credentials.

agent.ts
1import { Agent } from '@anvia/core'2import { LensClient } from '@anvia/lens'34const lens = new LensClient()5const tracing = lens.observer({ captureMode: 'safe' })67const agent = new Agent({8  id: 'support',9  model,10  observability: { observers: { tracing } },11})1213await agent.generate({ prompt: 'Summarize this ticket.' })14await lens.flush()15await lens.close()

Useful by default. Explicit when sensitive.

Safe capture exports operational metadata without input or output bodies. Full payloads remain opt-in; redaction and size limits are configurable and governed by your application.

Authorize + configure

Your runtime

Defines the agent, release policy, credentials, and which environments or data classes may include payloads.

observers: { tracing }

Isolated adapter

No global provider takeover. Payload capture is off until explicitly enabled.

metadata onlyredactionsize limits
project-scoped OTLP
  • Operational traces

    Agent, model, tool, environment, release, latency, usage, and errors.

  • Evaluation evidence

    Run lifecycle, cases, metrics, outcomes, usage, and direct trace references.

  • Storage you operate

    Store project-scoped traces and evaluation evidence in your Lens deployment.

Use flush() for an explicit delivery checkpoint; call close() during final cleanup.

Start with the signal. End with the cause.

Three repeatable paths from an operational or quality signal to the run-level evidence needed to act.

01 / LATENCY

Find a latency regression

Compare P95 by release and service, open the slow cohort, then inspect the model and tool spans consuming the budget.

  1. P95 alert
  2. release cohort
  3. slow trace
  4. span timing
release contexttrace timing
02 / TOOLS

Explain a tool-failure cluster

Group errors by service and tool, follow one representative run, and separate bad input from permissions, retries, or provider failure.

  1. error cluster
  2. tool span
  3. run context
  4. root cause
tool eventserror context
03 / QUALITY

Resolve a failed quality gate

Open the candidate comparison, inspect the failing cases, jump to their traces, then rerun the same published dataset version.

  1. blocked gate
  2. failed case
  3. linked trace
  4. rerun
managed datasetrelease gate

Debug locally. Learn in production.

Studio helps you understand an agent before it ships. Lens is an optional production companion that keeps the same runtime concepts visible after deployment. Your agent can also use another observability backend.

Same agent and trace concepts, from your machine to production.

Run Lens on your own infrastructure.

Download the Compose and environment templates, configure unique secrets, then start the stack. The first account becomes the workspace owner.

  1. 01

    Prepare the deployment

    Download the official templates, pin a Lens release, and configure unique secrets.

  2. 02

    Start the workspace

    Launch the stack. The first account created becomes the workspace owner.

  3. 03

    Connect a project

    Create a project and ingestion key, then send native Anvia telemetry or supported Langfuse v5 OTLP traces.

Make the evidence part of the release.

Connect one agent in safe-capture mode, investigate its traces, then promote a stable evaluation suite into an explicit quality gate.

Anvia