KazenAI Agent Lens

KazenAI Agent Lens product surface

After the run: timelines you can read, drift you can act on, and replay that does not oversell determinism.

Project information

  • Category: Agent observability · Drift · Post-run evidence
  • Focus: Timelines, release comparison, honest replay reporting
  • For: Teams past “it worked in the demo”
  • Role: Reference system behind reliability engagements

Overview

When an agent fails in production, teams need a record they can trust. Not a black box. Not marketing that pretends every “replay” re-runs the world identically.

Agent Lens gives operators readable run history, drift signals before the next promote, and analysis grounded in what was observed.

Work with me:

Ways I can help

Back to Projects

Problem

Teams often cannot explain agent runs, drift, or token spend after an incident or surprising output.

Mechanism

Trace timelines, drift checks, and replay analysis work from recorded evidence instead of silently re-running models.

What It Proves

Post-run review can diagnose behavior and cost without creating a second hidden model run.

Engagement Relevance

Useful for audits and retainers where operators need stable evidence for failure review and monitoring decisions.

CTA

Use this pattern to inspect whether current traces are enough to trust production agents.

Send a problem brief

Technical evidence

Frozen run trace and trace-grounded replay report (reference system example).

model_call   planner   Fix the auth token expiry bug in src/auth.py
tool_call    read_file src/auth.py
tool_call    read_file tests/test_auth.py
model_call   coder     Fix the expiry comparison to use UTC timestamps
tool_call    write_file src/auth.py

replay_mode:    trace_grounded_report
trace_digest:   sha256:8f3a…c21e

Known limitation: Default replay analyzes recorded evidence; bit-identical live LLM re-execution is a separate, gated mode and is not the default production story.

Goal

Make agent behavior inspectable after the fact and cautious before the next release. Timelines operators can read. Drift awareness. Honest labeling of what replay means.

What it does

  • Run timelines: a clear path through model, tool, and gate-style decisions for a given run.
  • Drift awareness: surfaces that help teams notice when behavior shifts across releases.
  • Replay reporting: analysis grounded in observed traces; live re-execution is a separate, gated mode when used at all.
  • Operator evidence: built for incident review and reliability conversations, not vanity metrics alone.

Honesty by design

Default replay is framed as analysis of what already happened. That keeps buyers and operators from being misled about determinism.

  • Trace-first storytelling: freeze and explain the observed path before claiming re-execution.
  • Cautious promotion signals: drift views prompt judgment; they do not manufacture false confidence.
  • Shared reliability narrative: cost blocks and human gates can appear in the same evidence trail as model steps.

Where it fits

  • Use cases: production agent fleets, pre-deploy checks, incident postmortems.
  • Signal: especially useful when demo success is no longer enough.

What I built

I built Agent Lens as an evidence product: visibility, drift caution, and honest replay framing. Observability that supports trust, not another opaque dashboard.

Why it matters for clients

Agent systems become operable when post-run evidence and release caution get the same seriousness as the agent itself. That is what audits surface and pilots install.

Send a problem brief   See engagement options