seekrit
← all posts

Deploying AI agents without a secrets plan is how keys leak

Agents

Most agent launch plans cover the model, the tools, the evals, the cost per run, and the guardrails on output. Credentials rarely get a line. During the first prototype the agent needed a Stripe key to test a tool, so the key went into a .env file or the env block of an MCP config. By launch the file holds eleven credentials, the agent runs with all of them in scope, and nobody can say which ones it uses.

Secrets management for a normal service is well enough understood that teams stop thinking about it, and from the outside an agent looks like a service: a process, some environment variables, an HTTP client. The differences are internal, and they are the reason the service habits don't transfer.

How an agent differs from a service

It takes instructions from its data. A service uses a credential the way its code says to. An agent uses it the way it was persuaded to, and the persuasion arrives through the same channel as the work: a web page it summarizes, a PR title, a tool result, a file in the repo. In one incident, a pull-request title was enough to make three coding agents post their own credentials to the thread.

There is no code path to review. With a service you can grep for where STRIPE_KEY is read and audit the call sites. An agent's call sites are chosen at runtime by a model. You cannot review where a credential in its environment will go; you can only put limits around it.

It holds many credentials at once. A service needs a database URL and a key or two. An agent with a dozen tools has a dozen keys in scope on every turn, whichever tool the turn uses. One bad turn exposes everything the agent can reach.

It gets copied. Agents run in sandboxes that snapshot the filesystem, in sessions that keep context, behind tracing that records every tool call, and in checkpoints that restore a rotated key months later. A credential in an agent's environment ends up in every trace, snapshot, and log line where someone serialized the environment to debug a run.

The .env model assumes plaintext in the environment is acceptable because only trusted, reviewed code reads it. An agent is neither, on purpose.

What goes wrong

None of these need a skilled attacker.

  1. Exfiltration by instruction. Content the agent reads tells it to send the key somewhere. It has the key and an HTTP client, so it does.

  2. Misuse of a legitimate key. The agent holds a real Stripe key to read charges, and one confused turn issues a refund, because the key permits both and nothing between the agent and Stripe distinguishes them.

  3. Sprawl. The key started in .env. It is now also in .mcp.json because the config format asks for it, in the trace store because the tool call was logged with headers, in the sandbox image because the build copied the file, and in a Slack DM because that was the quickest way to set a teammate up.

  4. One process boundary around every key. The Slack key and the database password share an environment. A vulnerable dependency in one tool (LangGraph had one) exposes all of them.

  5. No revocation story. Something looks wrong in the logs. Which agent? Which key? Each key is shared across every agent instance and probably CI. Rotating it means finding every copy from point 3, so it doesn't happen.

In each case the credential is plaintext, at rest, shared, and readable by a process whose behaviour cannot be enumerated.

Questions to answer before launch

  • Which credentials does this agent need, per tool?
  • Where does plaintext exist, and for how long? A file, the agent's context, one short-lived process, or nowhere on the agent's side?
  • What stops a valid credential from reaching the wrong destination, or the wrong operation on the right one?
  • Can the agent read or change the policy that constrains it?
  • When something goes wrong, what do you revoke, and what else breaks?
  • What is logged, and does anything fire only on misuse?

The rest of this post is how seekrit answers them, and where the answers stop.

Deploying an agent on seekrit

seekrit is a zero-knowledge secrets manager. Values are encrypted in your browser or CLI before upload, the server holds only ciphertext, and decryption happens wherever a service token is, never on our side.

One token per agent, bound to one environment

Each agent gets its own service token, bound to an environment that holds only the secrets that agent should reach. The token is the only credential in the agent's configuration. The MCP config, container env, or sandbox spec holds one revocable value instead of a dozen.

seekrit token create --app support-agent --env prod --name support-agent-prod

Revoking it affects one agent. Every read it made was authenticated by it, so "which key did it use" has one answer.

Plaintext in one process, or none

For an agent whose tools you trust, the first step is to stop putting values in files. seekrit run resolves the environment and hands it to a child process for the life of that process:

seekrit run -- python agent.py

No .env, nothing on disk. This is the same posture as any service, with the same limit: the process can read the values, so an agent in that process can be talked into reading them. For an agent that takes untrusted input, treat it as the minimum.

The stronger setup is one where the agent never holds the key. It holds a placeholder, {{seekrit:STRIPE_KEY}}, and seekrit-proxy runs beside it and substitutes the real value on the way out. The proxy decrypts once at startup, refuses to start if anything fails to decrypt, and scrubs any value the upstream echoes back. The agent's context window, traces, sandbox snapshot, and checkpoint all contain only the placeholder. Five levels walks the gradient between the two setups. For frameworks where a separate process is awkward, the same substitution ships in-process for LangGraph, CrewAI, Pydantic AI, Mastra, and the OpenAI Agents SDK, with the same default-deny allowlist and a weaker boundary, since it shares the agent's address space.

The allowlist bounds destinations and operations

The proxy is default-deny. Each rule names an upstream and the secrets allowed to travel to it. A placeholder pointed anywhere else is refused with a 403 and never forwarded, so "the agent was told to send the key to attacker.example" becomes a log line.

Rules can also restrict what the agent may do with a credential it is allowed to use:

[[route]]
prefix = "/stripe"
upstream = "https://api.stripe.com"
methods = ["GET"]
paths = ["/v1/charges", "/v1/charges/*"]
allow = ["STRIPE_KEY"]

The support agent can read charges with a key that, at Stripe, could also refund them. The refund request is refused before it leaves the machine, with a message naming the constraint that decided.

For operations no rule should decide in advance, the proxy can stop and ask. A matching request holds just before dispatch, the prompt names the credential about to travel, and no answer is a refusal. The trust ratchet covers the related case where a run that has read something sensitive should be able to do less afterwards: a declared read withdraws hosts and secrets from the rest of that run.

The agent cannot widen its own policy

Rules can live in the proxy's local file or be written in the dashboard as signed policy. Publishing requires a human session and is signed in the browser with your own key. Each proxy verifies the signature against thumbprints pinned in its local config. seekrit stores and serves the bundle but cannot forge one, and an agent holding an admin token cannot publish one.

Short-lived credentials instead of standing ones

Where the upstream supports it, the agent holds no long-lived credential. Temporary access mints a Postgres, MySQL, Redis, MongoDB, SSH, AWS, or GCP credential that expires after an hour, with the password generated on the agent's side so the control plane never sees it.

Logging and tripwires

Each substitution the proxy makes logs the secret names, method, path, and upstream host, never the values. Each denied resolve is written to the org's audit log. For copies you may have missed, a honey token is a decoy credential that unlocks nothing and emails you when anything tries to use it. Plant one in the old .env, the archived MCP config, or the sandbox image you meant to rebuild. Nothing legitimate holds one, so a trip means the place it sat in was read.

The agent can do the setup

seekrit speaks MCP through two servers, split along the zero-knowledge line: a hosted metadata server the agent can sign up to with no credential and no browser, and a local crypto server that handles anything touching a value, on the agent's own machine. One machine credential drives both. An agent can create its own org, environments, and tokens, and hand the org to a human when it is done. The hosted server registers no tool that returns a secret value.

Where the answers stop

The in-process shim keeps values out of the framework's config and traces but not out of reach of code running in the same process. The proxy is the stronger boundary.

The proxy bounds where a credential goes and which operations it authorizes. It does not judge what an agent does within a permitted operation. A rule that allows POST /v1/messages to Slack allows any message. The approval hold covers that, and it needs a human on the other end.

A service token is still a credential and a leaked one still needs revoking. It is one thing, bound to one environment, with a record of everything it read.

seekrit run on its own removes files from the picture and nothing else. If the agent reads untrusted input, put the proxy between it and the credential.

Start

npx plugins add seekritdev/agent-plugin
seekrit token create --app my-agent --env prod --name my-agent-prod
SEEKRIT_TOKEN=skt_… seekrit proxy run --host api.stripe.com=STRIPE_KEY

The agent proxy guide covers the config, the framework guides cover the in-process option, and the sandbox guides cover keeping the credential outside a hosted runtime.

If the agent gets a placeholder on day one, there is no .env to migrate away from at launch.