Let Authority Follow the Work: Tenuo for NVIDIA OpenShell
TL;DR: Today we’re releasing tenuo-openshell, an open-source integration that adds task-scoped authorization, approvals, safe delegation, and signed receipts to NVIDIA OpenShell. OpenShell decides what a sandbox can reach. Tenuo decides what the task running inside it may do there, and checks every tool call before OpenShell attaches credentials. No fork, and no changes to your agent code.
If you already run a local OpenShell gateway, the demo takes about five minutes and needs no model or API key. Install the CLI, then:
tenuo-openshell dev up # prints the gateway.toml block and a sandbox command # add the block, restart the gateway, run the printed sandbox command tenuo-openshell provision --dev --sandbox tenuo-demo --preset demo tenuo-openshell demo call read_logs service=payments environment=staging # allowed tenuo-openshell demo call read_logs service=identity environment=production # deniedThe full quickstart walks through registration, the approval flow, and cleanup. The rest of this post explains how it works.
If you’re running agents in OpenShell, you already have the hard part of containment covered. Your agent runs as an untrusted program in a sandbox, it can only reach the destinations you allow, and it never holds your provider credentials. It can do useful work without having the run of your host.
Then the agent picks up a task. Say it’s helping with an incident on the payments service. One task needs to read the logs. Another needs to restart the service, and you’d like a human to sign off on that first. Later, the agent hands part of the investigation to a specialist agent, which should only be able to read, and only for a few minutes.
Both tasks run in the same sandbox and call the same service. OpenShell’s job is to decide what the sandbox can reach, and it does that well. Which of these calls each task is allowed to make is a different question, and it changes with every task. That’s the question tenuo-openshell answers.
How OpenShell keeps an agent contained
An OpenShell deployment has three parts. When an operator creates a sandbox, the gateway provisions the other two:
- Gateway. The control plane. It coordinates sandboxes, policies, providers, and logs, and drives the infrastructure that places each sandbox on a runtime.
- Workload. Runs the agent’s code, treated as an untrusted program.
- Supervisor. Runs beside the workload and owns the mediated path from it to everything else.
OpenShell separates the untrusted workload from the mediated path to external services. Credentials enter that path only after the request passes the configured controls.
The sandbox policy works at several layers:
- Filesystem rules limit which paths the workload can touch.
- Process controls block privilege escalation and other dangerous system behavior.
- Network rules decide which destinations the sandbox can reach.
- Protocol rules constrain HTTP, WebSocket, and MCP traffic. For MCP, policy can limit an approved client to specific tools.
Credentials take a separate path. OpenShell keeps them outside the workload and injects them only into traffic headed for an authorized endpoint, so the agent can use a service without ever seeing a reusable secret.
The supervisor is where our integration plugs in. After OpenShell’s policy has admitted a request, but before credentials are attached, the supervisor knows which sandbox sent it, where it’s going, and exactly what it says. OpenShell’s supervisor middleware interface lets you run your own checks at that point without changing the agent or forking OpenShell.
What each task may do
OpenShell’s policy describes the sandbox, and a sandbox usually outlives any one task. You could rebuild its policy every time the task changes, but that pushes application logic into your runtime config. Or you could let every task do everything the sandbox can reach, which treats reachability as permission.
We’d rather keep the two separate:
| Layer | Question | Examples |
|---|---|---|
| OpenShell runtime policy | What may this sandbox reach? | Processes, filesystem paths, network destinations, credential use, MCP tool names |
| Tenuo task authority | What may this task do there? | Exact tools, argument constraints, expiry, approval, delegation chain |
Both have to say yes before a call goes through.
Tenuo carries the task’s authority in a warrant: a small signed token that says which tools the holder may use, which arguments it may send, how long it lasts, and how much of it may be passed on. An issuer you trust signs one per task, and tenuo-openshell provision installs it in the sandbox. In production, your service can issue one per request.
The integration ships with a small incident-operations example. Its MCP server exposes two tools, read_logs and restart_service, and OpenShell allows the sandbox to reach both. The task’s warrant is narrower. Here’s a readable view of it:
{
"capabilities": {
"read_logs": {
"service": "payments",
"environment": {"one_of": ["staging", "dev"]}
},
"restart_service": {
"service": "payments",
"environment": "staging",
"replicas": {"range": {"max": 5}}
}
},
"required_approvers": ["operator_public_key"],
"min_approvals": 1,
"approval_gates": {
"restart_service": "whole_tool"
}
}
The agent can read payments logs in staging or dev. It can propose restarting payments in staging with up to five replicas, but the gate on restart_service means an operator has to sign that exact call before it runs.
Inside the same sandbox, that gives you four different outcomes:
| Attempt | Outcome |
|---|---|
Read payments logs in staging |
Allowed |
Read identity logs in production |
Denied, because the arguments fall outside the task |
Restart payments with three replicas |
Held until an operator approves that exact call |
| Call the server without a warrant | Denied at the OpenShell middleware |
For the restart, the operator sees exactly what they’re approving:
tool restart_service
arguments {"environment":"staging","replicas":3,"service":"payments"}
warrant tnu_wrt_01a10adf5a937576b63088480b84ceb3
expires in 300 seconds
Their signature covers the warrant, the tool, the agent holding the warrant, and those arguments. It can’t be reused to restart eight replicas or a different service, and in production the replay store makes sure it’s used only once. The human gets pulled in for the one step that needs judgment, rather than watching every step or granting an open-ended admin session.
When the demo finishes, the MCP server’s log shows only the calls that were allowed. Denied calls never become authenticated upstream requests.
The same checks hold when a model chooses the calls. The NeMo Agent Toolkit guide runs a NAT ReAct agent in the sandbox, with NVIDIA’s API attached through an OpenShell provider profile, so the sandbox only ever sees a placeholder for your key.
We gave it the incident on nvidia/nemotron-3-super-120b-a12b. When its read of production logs was denied, it kept looking for a way in, retrying with identity-service, auth, auth-service, and prod for production. The warrant denied every attempt, and none reached the MCP server. That persistence is what makes an agent useful, and the warrant lets it keep going without giving it a way out of the task.
Checking the call before the credential goes on
The integration hooks into OpenShell’s documented supervisor middleware interface at HTTP_REQUEST / PRE_CREDENTIALS. Where that sits in the request path is the whole design:
The middleware makes the task-level authorization decision before OpenShell attaches the credential that lets the call become an authenticated upstream request.
Your agent’s MCP client talks to a small proxy on loopback inside the sandbox. The proxy holds a key generated for this task and signs each tools/call. The proof covers the warrant, the tool, and the arguments for a short time window, and the middleware accepts that signed call once unless the tool is marked safe to repeat. Your agent framework doesn’t need any Tenuo code.
OpenShell applies its own network and protocol policy first. For traffic it admits, the supervisor calls the Tenuo middleware, which:
- Authenticates the OpenShell caller and picks policy by the sandbox’s immutable
sandbox_id. - Confirms this sandbox may send the requested tool to this destination.
- Verifies the warrant chain back to a trusted issuer.
- Checks that the caller holds the private key the warrant names.
- Checks the tool and every argument against the task’s constraints.
- Checks approvals, replay, expiry, and revocation.
- Returns allow, or a denial with a reason code.
If the call is allowed, the middleware strips the Tenuo proof, and OpenShell injects the credential and sends the request upstream. Because the check runs before injection, the middleware never sees your provider credentials.
The check is local and fast. A decision takes hundreds of microseconds, including in production with TLS and a shared replay store, which is small next to a single model response. The benchmarks are in the repo.
The proxy also catches bad calls early, so the agent gets a useful error without the request leaving the sandbox. But the proxy shares the sandbox with the agent, so we don’t rely on it. If a compromised process skips the proxy, its call still lands at the middleware and gets denied before credentials go on. You can try this in the demo by sending a call with --unsigned.
Handing part of the job to another agent
Back to the specialist. Say the incident agent wants another agent to dig into a dependency. That agent might run in a different process, a different OpenShell sandbox, or a different framework entirely. Handing over the parent’s credential would hand over everything else it can do. Creating a standing role for the specialist would describe it in general, when what you want is permission for this one investigation.
Instead, the parent attenuates its warrant. The child warrant can drop tools, tighten argument constraints, shorten the lifetime, and block further delegation. It can’t add restart_service to a read-only parent, and it can’t outlive the warrant it came from.
The specialist generates its own key, and only its public key and the narrower warrant cross over. When it calls a tool through OpenShell, the middleware verifies the whole chain and the specialist’s signature. Each runtime keeps its own service credentials behind its own boundary. Only the warrant travels.
The demo does this across two sandboxes. The agent in the first delegates read_logs to a key generated in the second, and only public keys and the warrant cross between them. The second sandbox’s read is allowed. Revoke the first task’s warrant, and the child’s next read is denied, because the revoked warrant is still in its chain.
The same chain can be checked at several points along the way: in NVIDIA NeMo Agent Toolkit before a function runs, at the OpenShell boundary before credentials go on, at an A2A worker receiving the handoff, and at an MCP server just before the effect happens. Each one sees the same record of how the task’s authority got there, so you don’t have to rebuild the policy at every handoff.
Each checkpoint also signs a receipt for what it decided, with its own key: the OpenShell middleware, the NeMo Agent Toolkit plugin, and the MCP server. A receipt commits to the warrant chain, the holder’s proof, and the request, and each log is hash-chained, so deleting or reordering an entry breaks every link after it. The demo checks all three sets offline: every expected allow and denial was recorded where it should be, and delegated calls name both the parent and the child.
What this doesn’t cover yet
Today the integration authorizes MCP Streamable HTTP traffic crossing the configured middleware binding. Your OpenShell policy needs to close the paths it doesn’t inspect, including raw TCP, tls: skip endpoints, binary WebSocket frames, and server-to-client WebSocket messages.
A warrant limits what a task can do. It doesn’t tell you whether the model understood the incident correctly. If the model is wrong or manipulated, narrow arguments, short lifetimes, approvals, and revocation limit the damage.
Dev mode is built for a quick demo. Production needs what you’d expect on an authorization path: TLS, authenticated extension calls, signed policy, a shared replay store, and a plan for revocation and receipts. The production guide covers each of these, and the threat model lists the trusted components, attacker positions, and residual risks.
None of this needs a hosted service. The middleware reads local policy, verifies warrants locally, and keeps receipts on disk. When you want to manage many gateways from one place, a control plane can deliver signed policy and revocations and collect receipts through a provider contract that keeps it off the decision path: if it goes down, calls keep being decided on the last valid policy until that policy’s freshness deadline passes.
If you run agents in OpenShell, try the quickstart on one task and tell us what breaks. Issues are open.
Further reading
Result receipts. With result evaluation on, the middleware also signs a receipt for the result of each allowed call, with the result’s status, size, and digest, linked to the decision it follows. A decision receipt doesn’t claim the effect completed; a result receipt shows a result came back. The end-to-end demo verifies both offline.
Least privilege, per task. Saltzer and Schroeder wrote in 1975 that “every program and every user of the system should operate using the least set of privileges necessary to complete the job.” For an agent, the job can change with every request while the process and its identity stay the same. A stable sandbox plus a per-task warrant lets you honor both timescales, and gives you the confidence to let agents do more: investigate across systems, take bounded action, ask a human at the right moment, and delegate a smaller piece of the work.
Beyond OpenShell. Tenuo has integrations for OpenAI, Google ADK, LangChain, LangGraph, CrewAI, AutoGen, Temporal, MCP, A2A, FastAPI, and NVIDIA NeMo Agent Toolkit, so one warrant chain can cross a mixed agent system. The delegation model is written up in the Attenuating Authorization Tokens Internet-Draft, submitted to the IETF OAuth Working Group, with Tenuo as its reference implementation.