Observe: what your agents do and cost
Every agent session, coding or custom, in one place: the trajectory, the model, the tokens, the cost, and what kind of work it was.
You are paying for agents in two places: the coding agents your engineers run all day, and the agents you build into your product. Both produce the same thing, a session: a user asks, a model answers, tools run, something is committed or returned. Observe records that session once, in one shape, and answers the questions that are otherwise guesswork. What did this week’s agent work cost? How much of it went to fixing bugs versus shipping features? Which model is doing the work? What did the agent actually do in the run that failed?
Two ways in
The agents you run. Install the gjalla plugin in Claude Code, Codex or Hermes Agent and every session is recorded through the harness’s own hooks; nothing to run by hand. gjalla scan backfills the last 14 days from the transcripts already on your machine, so the first report is available in minutes, before the plugin has seen a single new session. Get started with the agents you run →
The agents you build. Add gjalla-observe to the service that runs your agent. One line for LangGraph and the OpenAI Agents SDK; one context manager around your job runner for a hand-rolled loop; a hook argument for Strands; a ten-line recipe for anything else. Get started with the agents you build →
Both land in the same table.
What you see
Observe › Sessions lists every session: harness, model, tokens, cost, commits, task type, when it started and ended. Open one and the trajectory is there in order, each user turn, each assistant turn with its model, each tool call with its input and each result with its size (and the error excerpt when it failed), plus the commits the agent made and the one-line attestation it wrote for each.
Measure is the rollup over the last 30 days: sessions, tokens, cost (with the share that was cache reads), how much went to bug-fixes, where tokens go by kind of work, and a by-model table.
How a session gets a kind of work
Coding agents classify their own commits. When an agent commits, the plugin records the commit against the session and hands the agent one pre-filled line to run:
gjalla attest add --commit <sha> --session <id> --task-type bug-fix --summary "..."
The task type is one of feature, bug-fix, refactor, docs, test, chore. A session’s type is the type recorded on its commits when there is one; full teams with transcript analysis on also get a judged type from the weekly analysis for sessions with no commit. Measure’s “where tokens go” is built from these.
Inside a typical session
Every judged session also gets a context breakdown: how many of its input tokens are the static prefix re-read on every turn (system prompt, tool schemas, MCP context, skills, agent definitions, hooks, and reminders) versus tokens spent making progress. It’s metadata only, token counts by kind, never the prompt or response text.
Per-session judgement also labels step ranges with a fixed taxonomy, so Measure can say not just how much was overhead but what kind and why:
- Kind:
productive,overhead,rediscovery,rework,dead_end - Cause:
implement,verify,investigate,plan,orientation,setup,tool_friction,context_load,known_code,known_fact,repeated_investigation,own_bug,misread_request,retry_loop,reverted,unverified_claim,abandoned_approach,off_task - Avoidable by:
memory,rule,skill,clearer_prompt,tooling,none
Only full teams with transcript analysis on get judgement; a session’s task type and cause labels come from that weekly run.
What is collected, and what is not
Collected: user and assistant text (capped at 32 KB per turn), the model and token usage per assistant turn, tool names and inputs (capped at 8 KB), tool result sizes and error classes with a 200-character excerpt, commits and their attestations, context compactions, interrupts, subagent boundaries, the repo remote and branch, and a hash of the working directory.
Never collected: tool result bodies, file contents, or anything about your machine beyond that hash.
Where sessions land
Sessions from your own machine go to your personal team by default. To pool a team’s sessions, each member runs gjalla auth team <team-id> once; sessions then land in that team’s Observe. A custom agent’s telemetry connection names its team and project directly.
