MCP setup
Install and configure the lucen-mcp server, every tool it exposes with the scope it needs, and which of them wait on a worker you start.
The hub is designed to be driven by an agent, not just by a human with a
browser. lucen-mcp is that surface: an MCP
server exposing the hub's tools over the same /v1 endpoints the web app and
the CLI use. This page is honest about which of them do something today, and
about which wait on a machine you start.
Install
There is no lucen-mcp on PyPI — if you find a package by that name, it is
not ours. This hub serves it:
curl -fsSL https://api.mousemouse.ai/install.sh | sh -s -- mcp
That puts a lucen-mcp executable on your PATH. It speaks MCP over stdio,
takes no arguments, and is configured entirely by environment variables.
Configure
Claude Code — .mcp.json at the root of the project you want it in:
{
"mcpServers": {
"lucen": {
"command": "lucen-mcp",
"env": {
"LUCEN_API_URL": "https://api.mousemouse.ai",
"LUCEN_API_KEY": "lucen_sk_..."
}
}
}
}
Claude Desktop — the same entry, in
~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or
%APPDATA%\Claude\claude_desktop_config.json (Windows). Claude Desktop does
not read your shell profile, so if lucen-mcp is not on the launcher's PATH,
use the absolute path uv tool install printed — usually
~/.local/bin/lucen-mcp.
LUCEN_API_URL is the hub's API — https://api.mousemouse.ai for the hub serving this page.
It is required: the server has no built-in hub and refuses to start without
it, rather than send a key to a guessed address. LUCEN_API_KEY is
optional: without it the server still runs and every none-scope tool
works against public data. Create a key in API keys.
Two rules that will not change:
- The key is an environment variable, never an argument. No tool takes a credential as a parameter — it would land in the model's transcript — and there is no command-line flag for one, which would put it in the process table and your shell history.
- Tools map to operations, one to one. There is no "do what I mean" tool: the contract in the API reference is the tool surface, so an agent's transcript stays auditable against it.
The tools
Scopes nest: read ⊂ write ⊂ train. none means no key is needed.
| Tool | Scope | What it does |
|---|---|---|
search_repos | none | Search datasets, models and robot repos |
get_repo | none | One repo with its dataset meta / model meta / robot card |
get_robot_card | none | DoF, mass, actuators, joint limits, mjcf_path |
get_dataset_episodes | none | The episodes in a LeRobot-format dataset repo |
search_experience | none | Search the 171 experience cards |
get_coach_doctrine | none | The 22 numbered rules a report cites by id |
get_training_advice | train | Diagnose a run — spends model budget |
get_training_plan | train | Plan a programme from a task brief — spends model budget |
get_run_status | read | One training run: status, recipe, lineage, artifacts |
get_run_logs | read | The last N lines of a run's log plus its metrics, replayed once (ANSI stripped by default) |
get_run_config | read | A run's resolved config grouped by RL structure, and what changed against its parent — readable while it trains |
get_run_gates | read | The sim gates a run's checkpoints got while it trained: step × cells × verdict, with posters |
render_rollout | train | Render a gate's three views (side · front · feet) now — seed 0 replayed with video, checked against the measured episode; unbilled CPU |
cancel_training | train | Cancel a run — queued ends now, running is stopped by its runner |
estimate_training | train | Price a recipe before submitting it — creates nothing |
recommend_training | read | Which GPU tier a recipe needs — the hub's deterministic planner shows its arithmetic; creates nothing |
submit_training | train | Queue a run for a self-hosted lucen worker — dry run by default |
rollout_in_sim | train | Queue a sim rollout for a rollout worker (lucen rollout worker) — the hub scores it PASS/FAIL per cell; dry run by default |
get_rollout_status | read | One rollout: status, claimer, per-cell verdicts, presigned video / trajectory / scorecard |
| request_deploy | write | Ask for a policy on a robot — a human must approve in the web app; starts nothing. With endpoint_id, ask for the robot to consume a hub endpoint instead |
| get_device_status | read | A device's pending requests, approvals and the driver's signed trail |
| serve_model | train | Create a hub inference endpoint for a skill-layer model; a reflex-class model is the hub's 409. Waits for a lucen serve you start; dry run by default. No tool opens a session |
| get_endpoint_status | read | One endpoint: status, tier, server, and per-session metering (chunks, GPU-seconds, service and round-trip p50/p95, cost) |
| plan_task | train | The hub's planner decomposes a task on a device into a closed step vocabulary and files the approval requests it needs. A human approves each handoff; the plan never approves, never opens a session, never names a joint. Costs money per decision |
| get_plan_status | read | One plan: status, steps, the timeline with each decision's cost, and what it is waiting on |
| create_repo | write | Create an empty dataset, model or robot repo — private unless told otherwise |
| validate_robot | read | Ask the hub to check a robot card against the repo's MJCF without writing: checks[] with a stable code, the offending names and a one-line fix. The call to iterate against |
| set_robot_card | write | Store a robot card (full replace). A card the hub refuses comes back as one line per failed check, and nothing is written |
| init_robot | local | Compile a local MJCF or URDF with MuJoCo and draft robot_card.json — card, interface, and a needs_you list of everything a model file cannot state. No hub call, no key |
| publish_robot | local, then write | Validate a local robot directory in MuJoCo (does it stand on its declared gains?) and publish it. Dry run by default: returns the report, uploads nothing |
| list_my_orgs | read | The organizations your key's user belongs to, with the role in each — the slug is the owner of anything you create for the lab |
| create_org | write | Create an organization; your key's user becomes its first owner. The slug shares one namespace with usernames (a taken one is the hub's 409) |
| add_org_member | write | Add a person to an organization by their hub handle, or change their role (owners only). There is no e-mail invitation: someone who has not signed up on this hub cannot be added, and the tool says so. Removing a member and leaving are a person's own, on /settings/organizations or lucen org |
The two coach searches are the reason to attach the server at all, and they cost nothing: 171 cards and 22 doctrine rules distilled from real robot-learning runs, each carrying the sentence from the war history it came from. Search them before spending a paid call.
Adding a robot
Five tools take a robot from a file on your disk to a robot page
(Add your robot is the same flow by hand). Two of them
are local: an MCP server runs on your machine, next to the model files, so
init_robot and publish_robot compile and simulate there. They need the CLI's
rollout extra importable in the server's environment, and say so — with the
install line — when it is not.
The loop an agent runs: init_robot → read needs_you, edit
robot_card.json → publish_robot (dry run) → read report.checks[], fix what
failed → publish_robot with dry_run: false. The hub rejects with reasons
and never repairs: a joint name is never corrected and a missing mesh is never
skipped. validate_robot returns the hub's own verdict on a candidate card
without writing anything, so a wrong card costs a read, not a refused write.
None of this needs train. Adding a robot moves data, never money.
What is not built, and how the tools say so
Three of the tools queue work the hub itself cannot execute, and the hosted paths are in the contract with no implementation behind them. An agent that calls a tool which silently does nothing is worse than one that gets told "not available yet", so:
submit_trainingdefaults todry_run: true— it returns the exact request body without sending it.POST /v1/runsrecords a run and hands it to an executor. The only executor this hub can drive isworker: alucen worker runprocess a human starts on their own GPU machine claims the run, executes its template, streams logs and publishes the model repo. The hub does not track connected runners, so the tool says the run is queued for the owner's runner, never that training started; a run nobody claims staysqueuedwith no weights and no bill. The hub's own GPU backend (modal) is real only on a deployment that names a deployed Modal app; elsewhere it answers 503 naming that setting and the tool relays it.rollout_in_simalso defaults todry_run: true.POST /v1/rolloutsrecords a battery of command cells asqueued; a rollout worker (lucen rollout worker, the CLI installed aslucen-cli[rollout], started by a human on their own CPU machine) claims it, runs the policy in MuJoCo on the robot's own MJCF under the policy'sio_contract(a mismatch is refused, never run), and reports measurements — the hub scores them PASS/FAIL per cell and, when the rollout names arun_id, attaches the scorecard to that run so the Coach reads every PASS cell as a standing constraint. The hub does not track workers, so the tool says the rollout is queued for a worker, never that it is running;get_rollout_statusshowsclaimed_byonce one has it and the verdicts once it is scored. Withexecutor: "modal"the hub runs it on its deployed Modal app instead — only where the deployment names one; elsewhere that is a 503. For a robot with no policy at all, the Sim tab on its repo page still runs the MJCF interactively in the browser.serve_modeldefaults todry_run: truetoo, and returns the estimate with the body.POST /v1/endpointsrecords an inference endpoint for a skill-layer model; the hub does not run it. Withexecutor: worker(the default) it waitsqueueduntil a human startslucen serveon their own GPU machine, which speaks openpi's policy-server protocol; the hub's own GPU backend (modal) spawns a container that claims it where the deployment names a Modal app, and answers 503 elsewhere. A reflex-class model — a 50 Hz locomotion or balance policy — is refused with 409: it runs on the robot, never over a network. No tool opens a session on an endpoint; a robot reaches one only through an approval a person signs (request_deploywithendpoint_id), and holds on its own clock when the link degrades.
get_run_logs reads the run's real log stream, replayed once (a tool call has
to end): the last tail lines, every metrics line, and the run's status when
the stream closed. A notice the hub wrote into the log — a cancel request, a
lease re-queue after the runner stopped reporting — comes back as stderr,
so an agent can tell the runner's words from the hub's. cancel_training
says whether the cancel was immediate (the run was queued) or a request the
runner honours on its next report (it was running; poll get_run_status).
The loop that works today is: train on your own hardware → record it with
lucen runs import → ask the Coach about it.
One deliberate absence, and one consent-gated handoff
There is no key management. /v1/keys is session-scoped in the contract:
no API key of any scope may mint, list or revoke a key. That is what stops a
leaked write key from promoting itself to train, and exposing those
operations here would expose something this server's credential can never do.
There is no tool that starts, stops or steers a robot — but there is a way to
ask. The device consent layer keeps one rule: approving a policy onto a
robot is session-scoped, so no API key can do it, including the one in
this server's env. What the agent gets is
request_deploy, which files an approval request, and get_device_status,
which watches what happens to it.
The flow, end to end
On the robot's own computer, install the driver, pull the policy whose contract this robot runs, and register the device. Registration generates a keypair on the box, pins the hub's signing key, and tells the hub which IO contracts the device will accept — here, the contract stamped inside a real policy:
curl -fsSL https://api.mousemouse.ai/install.sh | sh -s -- device
device installs lucen beside lucen-device and points both at this hub.
lucen pull lucen/laika-omni-c4-ff800 ./omni
lucen-device register --robot lucen/laika --name lab-laika --adapter virtual --accept-from ./omni/omni_c4_ff800.onnx
The virtual adapter validates the policy and runs a clock at its control rate
— no physics, no inference, no actuators — so the whole path can be rehearsed
without hardware. command shells out to the real robot's own start/stop
commands, taken from the driver's local config and never from the hub. Any other
robot gets its own adapter as a plugin: Connect your robot.
From the agent: request_deploy(model="lucen/laika-omni-c4-ff800", device="lab-laika", minutes=2). The hub pins the request to the repo's ONNX
and its sha256 right now, checks the policy's io_contract.json against what
the device accepts, and records it as pending. The tool answers
human_must_approve: true.
A person opens the request in the web app's Inbox, reads the
hub's checks, and holds Approve (or denies it with a reason). That mints a
short-lived EdDSA token bound to the device, the digest and the window. Only a signed-in session can do this step — a train key gets
403.
On the robot, lucen-device serve (or --once) picks the approval up,
verifies the token offline against the pinned key, pulls the repo with every
digest checked, compares the ONNX to the digest inside the token, reads the
contract out of the ONNX itself and refuses a mismatch, starts the adapter,
heartbeats, and stops when the window ends:
lucen-device serve --once
Back at the agent, get_device_status(device="lab-laika") shows started,
the heartbeats and stopped — or refused with the reason. Until a started
event exists, the policy is not running, whatever the request said.
Goal commands and teleoperation through an agent are not built (v1+) and the server's instructions say so.
Without the MCP server
MCP tools are a thin shell over /v1. An agent can work against the hub with a
scoped key and plain HTTP, and always could:
curl -s "https://api.mousemouse.ai/v1/repos?kind=dataset&limit=5"
lucen repo ls --kind dataset --json
from lucen_sdk import Client # generated from docs/openapi.yaml by `make sdk`
Errors are RFC 9457 problem+json
everywhere, which is the difference between an agent that can retry correctly
and one that regex-matches a string. The MCP server translates them further: a
403 comes back naming the scope your key is missing, and a 404 is always "not
found, or not visible to this credential" — in the same words for a private
resource and an absent one, because a 403 would confirm it exists.
There is also llms.txt, a machine-readable summary of all of
this at the site root.
Catching up: the activity feed and the overview
Two read-only tools let an agent see what its person would see on the Inbox. Neither can approve, deny, start or stop anything.
| Tool | Scope | What it does |
|---|---|---|
list_activity | read | The hub's audit log as one sentence per row, newest first, limited to what your key's user could already see. actor.kind is person, key or system — a key is reported as a key, whoever held it, this server included — and actor.agent: planner marks the only rows where the hub knows an agent decided. Filter by actor, resource type, time or owner |
get_campaign | read | A campaign — a task on a robot with a budget and rounds: each round's one declared change, verdict, the checkpoint it resumed from, any override a person recorded, its cost, and what waits on a person |
get_round | read | One round: its locked spec, the verdict at a checkpoint (its LOCKED gates, the hub's band rule beside them as a measurement), how many resolved-config keys moved against the parent, and its decisions |
propose_round | train | Write the next round's spec and start its run — dry run by default. The hub refuses a round that breaks one of the three rules (one variable, resume only from PASS, a FAIL ends the lineage) or could overspend the budget, with reasons; no key can override a rule |
decide_round | train / write | iterate from a PASS checkpoint, or stop a lineage. Shortlisting for the real robot and closing a campaign are a person's decisions: this tool refuses both |
list_trials | read | Real-robot trials the key can see — a shortlist becomes a session of trials on one device, the control first and last — with each trial's status (needs_verdict waits on a person), tilt and criteria in a brief |
get_trial | read | One trial as a record: the run sheet a person approved, the approval it rode on, the log and summary the robot's driver uploaded signed with the device key, the criteria judged against what was written down before the run, the operator's notes, the A/B/A bracket. The verdict (promote / iterate / stop) is a person's; no tool sets it |
get_my_overview | read | What waits on a person and what is running, in one call: pending approval requests (the newest with the hub's checks), queued and running runs and endpoints, active plans, devices online, and today's model spend with the sentence that says how it was counted |
get_device_status now carries, for each pending request, what the hub checked
before a person decides: whether the device accepts the policy's contract, its
control class and where that class lets it run, the fail-safe probe the device
declared, the latest sim-gate scorecard for that model on that robot (or that
there is none), the policy's embodiment fingerprint against the robot's
interface, and whether the pinned bytes are still the repo's. Relay them to the
person; they inform a decision no API key can make.