Observe the agents you build
Add gjalla-observe to the service that runs your agent. LangGraph, LangChain, OpenAI Agents SDK, Strands, Claude Agent SDK, or your own loop; sessions land next to your coding agents.
The agents you build produce the same trajectory a coding agent does: a user turn, model calls with token usage, tool calls, a result. gjalla-observe sends that trajectory to the same sessions table, so a product agent’s runs sit next to your engineers’ Claude Code sessions with the same cost rollup. It is a thin Python SDK that speaks gjalla’s event schema directly: no OpenTelemetry setup, one dependency (httpx), Python 3.9+, and it never raises into your agent.
Let your coding agent do it
The gjalla-instrument skill (npx skills add gjalla/engineering) is this page in a form your agent can apply to your repo: it reads the code, finds where a session begins and ends, and adds the right block below.
1. Get a key
In the gjalla web app: Settings → Team Settings → Agents and integrations → Cloud telemetry connections → Connect provider, source Application SDK, pick the project the agent belongs to. Put the key in the service’s environment:
export GJALLA_API_KEY=...
That key names the team and project, so every session from the service lands in the right place without any header or config. (For a quick local try, gjalla auth login works too; those sessions go to your personal team.)
2. Install
pip install gjalla-observe
3. Add the block that matches your agent
The one decision is where a session starts and ends in your app: a job, a request, a conversation, a graph run. Each block below tells gjalla that boundary; the model and tool calls inside it are captured automatically.
LangGraph, or LangChain with a root run
import gjalla_observe
gjalla_observe.init(service="my-service")
One line at process start. Every run through LangChain is recorded; the session id is LangGraph’s thread_id when present, else the root run id, and the session ends when the root run ends.
A hand-rolled loop over LangChain chat models
This is the common shape in production: your own orchestrator calling ChatAnthropic or ChatOpenAI directly, no graph. init() alone would make one session per model call, so add a scope around the unit of work that is a session (the job runner, the request handler), not around the model calls:
import gjalla_observe
gjalla_observe.init(service="my-service")
async def run_job(job_id, payload):
async with gjalla_observe.session(str(job_id), harness="custom", service="my-service"):
return await execute(payload)
session() works as with or async with, and ends the session on exit (failed if an exception propagates). This is the only call-site change.
OpenAI Agents SDK
import gjalla_observe
gjalla_observe.init(service="my-service")
init() registers a trace processor. A trace is a session; give runs a group_id and the whole conversation is one session.
Strands Agents
from strands import Agent
from gjalla_observe import GjallaHooks
agent = Agent(hooks=[GjallaHooks()])
Claude Agent SDK
For production runs with no local transcript (a service, a worker, a CI job). In development, Claude Code already writes the transcript locally and the gjalla plugin records it, so do not wrap the SDK there.
async for message in gjalla_observe.observe_query(prompt="...", service="my-service"):
...
Anthropic or OpenAI SDK directly, or any other loop
from gjalla_observe import Session
session = Session(run_id, harness="custom", service="my-service")
session.user(user_message)
response = client.messages.create(...)
session.assistant(response.content[0].text, model=response.model,
usage={"input": response.usage.input_tokens, "output": response.usage.output_tokens})
session.tool_use(call_id, tool_name, tool_input)
session.tool_result(call_id, size=len(result), is_error=False)
session.end()
4. Run it once and look
Trigger one real run. On Observe › Sessions a row appears with harness custom (or your framework’s name), the model, and token totals; open it for the trajectory, every model call and tool call in order with inputs and error excerpts. On Measure the run is in the cost tile. That is the whole verification.
If the row does not appear, the SDK logged exactly one warning under the gjalla_observe logger with the server’s response. A 401 is a bad or revoked key; a 403 is a key that cannot write sessions (a connection whose source is not Application SDK).
What is sent
User and assistant text (capped at 32 KB), model and token usage per assistant turn, tool names and inputs (capped at 8 KB), tool result sizes and error classes with a 200-character excerpt, session start and end, and the repo remote and branch of the working directory when there is one. Never tool result bodies or file contents.
Prompt capture is a policy you control: init(capture_prompts=False) stops the SDK recording the input side of model calls, for teams whose prompts are their end users’ data. Session and session() take the same flag.
Test suites
Set GJALLA_OBSERVE_DISABLED=1 in your test environment. Every SDK call becomes a silent no-op, so a developer machine with a live key in ~/.gjalla/config.yaml never posts sessions from the test suite.
Other frameworks
TypeScript (@gjalla/observe, with Vercel AI SDK agents and Mastra first) is next. An agent already on OpenTelemetry can point its OTLP exporter at /api/otlp/v1/traces with the key as bearer token and gjalla.session.id on the root span, though that path does not get the trajectory view above. Ask for the framework you use: support@gjalla.io.
