ClawMetry

ClawMetry

Sign in to open the dashboard

Advanced: sign in with a gateway token
Initializing ClawMetry
Loading model, tasks, system health, and live streams…
Loading overview + model context
Loading active tasks
Loading system health
Connecting live streams
Can’t reach the collector on this machine, so some screens may look emptier than they are. Nothing has been lost — retrying.
🚀
Your 7-day trial has ended. Keep every runtime with a Pro license. Get a license I have a license key
Setting up your node. First check-in usually arrives in about 30 seconds.
💔 Check agent →
View Session
⚠️ See spending →
Today Loading task outcomes...
show details
🎯 How independent is your agent?
--
--
Last 7 days
No data yet
Session quality
--
🔬 Evaluators on your agent
Open Evals
⚖️ Did the change help?
Pick a comparison. One word, with the number of sessions it rests on.
Looking for recent changes...
⚙️ Advanced: compare two runs by id
Paste session ids; green = improvement, red = regression
🪥 Error triage
Mute known errors so they stop inflating counts
LIVE Loading...
auto-refreshes every 30s
Loading flow...
🏥 System Health
Services
Channels
Disk Usage
Cron Jobs
Sub-Agents (24h)
Heartbeat
🔍 Configuration Diagnostics
Loading diagnostics...
🐝 Active Tasks
⟳ 30s
🐝
Loading tasks...
🧠 Claude Opus ...
Waiting for activity...
-- -- --
--
--
--
❤️ Is your agent alive? ...
waiting...
Last check-in: --
--
--
Recent check-ins
--

Confirm

Guard

What is running right now, whether it has gone off track, and the button to stop it.

Running sessions

Loading sessions...

Policies

A policy watches for a detector signal and acts without anyone present. New policies start in monitor mode: they record what they would have done and change nothing. A policy can escalate over time, for example pause now and kill five minutes later if the agent is still stuck. Each step runs only if the agent is still flagged when its wait is up.

Loading policies...

Does it give the same answer twice?

An agent can go wrong by answering the same request differently each time. ClawMetry can measure this by replaying failed sessions, but only when you turn it on: each replay runs your agent and costs money.

Recent decisions

Loading...

Agent reported

Notes agents sent to their operators through the ClawMetry MCP tool, counted per category, and how often a detector finding was also reported by the agent. Uncorroborated means no independent evidence was found, which is not the same as false. See the behaviour signals these sit next to

Loading...

Signals

What people say to your agents and what the agents say back. Frustration, praise, refusals, work handed back, giving up, and retries, counted from the transcripts you already have. No model reads them.

Loading signals...

Each signal this window

Loading signals...

What each runtime exposes

A runtime that does not write user prompts to disk cannot have a frustration rate. That reads as "not exposed", never as zero. Presets are English first.

Loading signals...

Briefs

Save a question with a schedule and a channel. The answer arrives as a message, so you read it instead of asking it. Briefs are off until you switch one on.

Loading signals...

Alerts Pro

Get notified when something goes wrong with your agents.
Loading alerts…

Always on

These run without any rules from you. They're what raised the red banners at the top of the dashboard. Each one delivers in-app, plus any channel you've connected in Notifications. Use the switch to mute one, or click a channel pill to change where it goes.
Loading monitors…

Recent alert history

Loading history…
Delivery channels: Slack · Email · PagerDuty · Telegram shared with Approvals
Quality this week
Loading…

Loading grade…

Reading your agent's recent work.

What went wrong

Ranked by what it cost you.

    The rough runs

    Click any to see the trace, or turn one into a check.

      Grading uses signals we already track — tool failures, repeated calls, and edits that were never checked. Every flag shows its evidence.
      Add a judge key to also grade the writing quality of the answer.
      Grade the writing too →

      Already using an evaluation platform? Point it at this endpoint to pull every run with its outcome, cost and tokens: /api/otel/export?shape=sessions&window=7d

      ℹ️ Numbers below are for this computer only.
      📊 Today
      --
      📅 This Week
      --
      📆 This Month
      --
      📊 Token Usage (14 days)
      Loading...
      💰 Cost Breakdown
      Loading...
      🏆 Top Sessions by Cost which session burned the budget · top 20
      Loading...
      🧩 Cost By Plugin / Skill
      Loading...
      💸 Top Sessions by Cost Alert threshold: $ per session
      Loading...
      💱 Cost Comparison what same workload costs elsewhere · 30 days
      💡 Spend Optimization model downgrade suggestions · 30 days
      🔮 Trace Clusters auto-group sessions by behavior pattern
      Loading...
      📅 Activity Heatmap hourly usage intensity
      Loading...
      🎯 Skill Cost Leaderboard heuristic attribution, 30s window
      Loading...
      Loading...
      ⚠️ Anomalies detected in cron jobs, review health table below
      📊 Cron Health Monitor Click row to expand run history
      Loading health data...

      Edit Cron Job

      Memory
      Markdown files your agent reads and updates. Edit them too if you want.
      Summary
      All files
      Access log
      Explorer
      Loading files…
      Markdown
      Loading...
      Opening the trail…
      1

      What it was asked

      The request that started this session, and what the agent knew going in.
      The request
      Loading...
      What the agent knew
      Loading...
      2

      What it did

      Every step, in order. Replay the conversation, or open the timing and trace views.
      3

      How it ended

      Did the result match the request, and what did it cost.
      Verdict
      Loading...
      Quality score
      Loading...
      Spend and time
      Loading...
      Code it produced
      Loading...

      📈 Upgrade Impact

      Loading...
      ℹ️ Local node only. These figures reflect only this node's DuckDB. See the Fleet tab for cross-node totals.

      📊 History

      Token Usage Over Time

      Cost Over Time

      Active Sessions

      Cron Runs

      Loading...

      Snapshot

      
          
      Messages / min0
      Actions Taken0
      Active Tools0
      Tokens Used -
      How your messages get answered
      🧑 You send a message
      📨 Channels 0/min
      🔀 Gateway routing
      🧠 Brain model
      🛠️ Tools idle
      ↩️ Reply replies back to you
      replies back to you You 📱 TG 📡 Signal 💬 WA 🔀 Gateway Agent Runtime 🧠 AI Model unknown Auth: unknown 📋 Sessions ⚡ Exec 🌐 Web 🔍 Search 📅 Cron 🗣️ TTS 💾 Memory 🧬 Skills 💰 Cost Optimizer 🧠 Automation Advisor I N F R A S T R U C T U R E ⚙️ Runtime Node.js - Linux 🖥️ Machine Host 💿 Storage Disk 🌐 Network LAN 📨 Channels ➡️ 🔀 Gateway ➡️ 🧠 AI Brain ➡️ 🛠️ Tools messages in routes to AI uses tools
      📡 Active Session Lanes
      Loading…
      🔧 Live Tool Call Stream
      0 events
      Waiting for activity...
      🧠 Brain · Unified Activity Stream
      Activity density · last 60 min (30s buckets)
      Live event stream (newest first) new events
      Loading recent activity…
      🔄 Self-Evolve
      How your agent could improve.
      Structured review of recent activity. Each finding has evidence and a concrete action.
      Self-Evolve needs an Anthropic credential.
      Run claude once to sign in with OAuth (free, recommended), or export ANTHROPIC_API_KEY in your shell. We auto-detect, no copy/paste here.
      Where we looked
        📬 Notifications Pro
        One place to manage every delivery destination, used by Alerts and Approvals.
        Tracing
        Every session as a trace. Open one for the span tree, waterfall, and agent graph - with full inputs/outputs per span.
        Loading traces…
        Agent Graph
        Who spawned whom — cross-session agent topology from span data.
        Loading…
        Turn anatomy
        Decompose one agent turn into a waterfall: prompt → model call(s) → each tool (start→end) → compaction → reply. Bar width is proportional to wall-clock duration.
        Loading sessions…
        Tool catalog
        Every tool the agent actually invoked, grouped by provenance (builtin / MCP / plugin), with call count, p50/p95 latency and error rate. Click a tool to expand its recent calls.
        Loading tool catalog…

        🧠 Context usage

        How full your agents’ context windows get, when your agent compacted (proactively vs forced by an overflow), how many tokens each compaction reclaimed, and which sessions keep slamming into the wall.

        Loading utilization…
        Harness
        Everything wrapped around the model that turns "it can talk" into "it can work".
        What is a harness?
        An AI model on its own can only talk: you ask, it answers, it stops. The harness is all the machinery wrapped around the model that lets it actually do work: tools, memory, safety rails, and the loop that keeps it going until the job is done. Claude Code, Cursor, and your agent runtimes are all harnesses around the same few models. ClawMetry is a window into that machinery. Each part below links to the tab where you can watch it live. (Good primer: What's harness engineering?)
        🔁 The loop
        The agent reads, decides, acts, checks its own work, and goes again until the task is done. This loop is what makes it an agent instead of a chatbot.
        Watch it live in Activity →
        🔧 Tools
        Tools let the agent touch the world: read and edit files, run code, search the web, call other services. No tools, no work.
        See its tools →
        📚 Memory
        Models forget everything between sessions. The harness writes notes to files so your agent remembers you, your projects, and its own lessons.
        Read its memory →
        🪟 Context
        The model can only "see" a limited window of text at once. The harness keeps that window focused, trimming old noise so quality doesn't rot.
        See context usage →
        📦 Sandbox
        A safe room where the agent can run code and fail freely, without reaching things it shouldn't touch.
        Check its sandbox →
        🚦 Guardrails
        The rules for what the agent may do on its own and what needs your OK first, like deleting files or spending money.
        Review approvals →
        🤝 Teamwork
        For big jobs, one agent leads and hands pieces to helper agents, then stitches the results together.
        See who spawned whom →
        💬 Interfaces
        The same agent can meet you in a terminal, on Telegram or Slack, or in a browser. Different doors, one harness behind them.
        Follow the message flow →
        Is this repo ready for an agent?
        Agents get stuck on repos that do not explain themselves. This is what yours tells them.
        Scanning the repo…
        What this runtime uniquely exposes, beyond the generic tabs:
        Loading harness view…
        Logs
        Live log stream ● Connecting…
        Loading…

        🏊 Swimlane Compare

        Loading swimlanes...
        🛡️ Security
        ?
        Security Posture
        Scanning configuration...
        -
        Passed
        -
        Warnings
        -
        Failed
        🔒
        Tamper-evident log
        Checking the activity log for tampering...
        🗄️
        Data retention
        Checking how long this machine keeps activity history...
        days
        Threat Detection & Anomaly Alerts
        0
        Critical
        0
        High
        0
        Medium
        0
        Clean Sessions
        Threat timeline (newest first)
        Scanning...
        🗂️ Recorded findings
        Kept on disk, so findings stay here after the scan above moves on. Includes anything a connected security tool reported. Open a row to read the session it came from.
        Loading findings...
        🔒 Policy Events (PII · Injection · Credential Leak)
        -
        PII
        -
        Injection
        -
        Cred Leak
        Scanning...
        🔑 API Key Scan scanning...
        Scanning session events for leaked API keys...
        📋 Signature Catalog
        📋 Recent activity (approvals, budgets, pauses)
        No recorded activity yet. Approval decisions, budget changes, and pauses appear here.

        🔒 Tool Policy & Sandbox

        Loading tool policy…
        Loading approval audit…
        Approvals
        Pending Approvals
        Loading...
        Recent Decisions
        No history yet.
        Protection Rules
        Where approvals go
        Loading...
        🤖 Primary Model
        --
        🔄 Model Diversity
        --
        distinct models used
        Fallback Rate
        --
        💬 Total Turns
        --
        assistant responses tracked
        🤖 Model Mix
        Loading...
        📊 Per-Session Breakdown
        Model Sessions Turns Share
        Loading...
        🔒

        NemoClaw governance is part of ClawMetry Pro

        Pro adds the NemoClaw sandbox, policy drift detection, and one-click egress approvals so you can prove your agents are inside the lines.

        No credit card required. Cancel anytime.
        🆕 Skills
        Shortcuts your agent can use. Click a skill to browse its files.
        Loading...
        ClawMetry Cloud

        Monitor your agents from anywhere. Encrypted, real-time, accessible from any browser or as a Mac app.

        or
        or run clawmetry connect
        Cloud sync key

        This key decrypts your synced data in the browser. It never leaves this machine except end-to-end encrypted, and it is never shown on the hosted cloud dashboard.

        Self-Host

        Everything stays on this machine. Sign in once for a free 7-day Pro trial that unlocks every runtime.

        or
        I have a license key · or run clawmetry login

        💰 Budget & Alerts

        Current Spending
        Loading...
        Component
        🕰️
        ×
        Loading...
        Loading...