LiteLLM
A LiteLLM proxy is where an organisation's model credentials collect. One
process ends up holding every provider key in model_list, plus its own
DATABASE_URL, LITELLM_MASTER_KEY, and LITELLM_SALT_KEY. That makes it the
place where "which keys does this box have" matters most, and the place where a
.env file is least defensible.
There are three shapes, and they are ordered by how much the gateway holds.
Start here if seekrit is new: three
commands put the keys in
an environment, mint a token bound to it, and export SEEKRIT_TOKEN. A token
reads everything in its
environment, so
there is no per-key setup to do before any of the below.
1. Wrap the process
Nothing about config.yaml changes: LiteLLM resolves os.environ/NAME from the
process environment, and seekrit run is what fills it.
seekrit run -- litellm --config config.yaml
model_list:
- model_name: gpt-5.6-terra
litellm_params:
model: openai/gpt-5.6-terra
api_key: os.environ/OPENAI_API_KEY
Decryption happens in the short-lived CLI, which then execs the gateway. In a
container, seekrit-run is the same thing as a static
binary you put in front of the entrypoint.
This is the whole answer for most deployments. The next section is worth it when you want the gateway to pick up a rotated key without a restart, or when the platform gives you no way to wrap the process.
2. Read through LiteLLM's secret manager
LiteLLM has a first-class seam for this: a custom secret manager is asked
for every os.environ/NAME in the config, before the process environment is
consulted. seekrit ships one in the Python SDK.
pip install seekrit
Install it into the same environment as the gateway. There is no extra to add —
seekrit.litellm imports LiteLLM lazily, and it only ever runs inside a process
that already is LiteLLM.
The shim
LiteLLM loads a custom secret manager from a file beside config.yaml, not
from an installed package. So a two-line file is required:
# seekrit_secret_manager.py
from seekrit.litellm import SeekritSecretManager # noqa: F401
general_settings:
key_management_system: custom
key_management_settings:
custom_secret_manager: seekrit_secret_manager.SeekritSecretManager
access_mode: read_only
model_list:
- model_name: gpt-5.6-terra
litellm_params:
model: openai/gpt-5.6-terra
api_key: os.environ/OPENAI_API_KEY
Then run the gateway with a service token and nothing else:
SEEKRIT_TOKEN=skt_… litellm --config config.yaml
Write the path with exactly one dot. LiteLLM splits
custom_secret_manager on . into a file name and a class name, and imports
the file from the directory holding config.yaml. seekrit.litellm.SeekritSecretManager
fails at startup with too many values to unpack, and any dotted package path
would look for a file rather than the installed module. The shim is not
optional.
What it does that the interface doesn't imply
It caches. LiteLLM's get_secret has no cache of its own and consults the
manager on every lookup — at startup, on router rebuilds, and on the invoke
path. One seekrit resolve returns the whole environment, so the manager holds a
snapshot with a single TTL (five minutes by default) rather than fetching per
name. That TTL is also what bounds how long a revoked secret keeps working in a
running gateway.
A name it doesn't hold falls through. LiteLLM's contract for a miss is to
fall back to the process environment, which is what lets adoption be partial:
OPENAI_API_KEY from seekrit, DATABASE_URL from your deployment. Nothing has
to move at once.
A miss is loud. LiteLLM turns it into ValueError: No secret found in Custom Secret Manager for <name>, catches it, and logs it at ERROR with a
traceback before falling back. Those lines are the fallback working, not a
fault — the way to quiet them is to move the remaining names into the seekrit
environment.
A failure does not become a silent fallback. If the API is unreachable, the last good snapshot keeps answering for up to 24 hours and every serve of it is logged, rather than the gateway quietly picking up whatever its environment happens to hold. Past that bound the snapshot is dropped and the manager answers nothing.
It never writes. store_virtual_keys would have LiteLLM write its generated
virtual keys back through the manager; the seekrit SDK is read-path only, so
those calls refuse with an explanation. Keep access_mode: read_only and let
LiteLLM's virtual keys live in its own database.
Configuring it
key_management_settings has a fixed set of fields and drops anything else, so
the manager is configured two other ways. Environment variables:
| Variable | Default | Meaning |
|---|---|---|
SEEKRIT_TOKEN | The service token. Required. | |
SEEKRIT_API_URL | https://api.seekrit.dev | API base URL. |
SEEKRIT_LITELLM_TOKEN_ENV | SEEKRIT_TOKEN | Read the token from a different variable. |
SEEKRIT_LITELLM_CACHE_TTL | 300 | Seconds a snapshot is served before refetching. |
SEEKRIT_LITELLM_STALE_TTL | 86400 | Seconds the last good snapshot answers while refreshes fail. |
SEEKRIT_LITELLM_ALLOW | Comma-separated names this manager may answer. Everything else is a miss. |
Or subclass in the shim, which is the same file LiteLLM already loads:
# seekrit_secret_manager.py
from seekrit.litellm import SecretResolver
from seekrit.litellm import SeekritSecretManager as _Base
class SeekritSecretManager(_Base):
def __init__(self):
super().__init__(
resolver=SecretResolver(
cache_ttl=60,
allow=["OPENAI_API_KEY", "ANTHROPIC_API_KEY"],
)
)
allow is a courtesy, not a boundary — the token still resolves the whole
environment into this process. To narrow what the gateway can read, mint a
token bound to a narrower environment.
3. Never hold the provider key
Both sections above end with the plaintext in the gateway's memory for as long as it runs. For a multi-tenant gateway that is the thing worth removing: put a placeholder in the environment and let the egress proxy substitute the real value on the way out.
seekrit proxy init --mode forward \
--preset openai --preset anthropic --preset gemini
One preset per provider the gateway routes to, in forward mode. Point the gateway at it:
export HTTPS_PROXY=http://127.0.0.1:8081
export SSL_CERT_FILE=$PWD/seekrit-proxy-ca.pem
export REQUESTS_CA_BUNDLE=$PWD/seekrit-proxy-ca.pem
export OPENAI_API_KEY='{{seekrit:OPENAI_API_KEY}}'
export ANTHROPIC_API_KEY='{{seekrit:ANTHROPIC_API_KEY}}'
export GEMINI_API_KEY='{{seekrit:GEMINI_API_KEY}}'
litellm --config config.yaml
config.yaml is unchanged — api_key: os.environ/OPENAI_API_KEY still reads
the environment, and what it finds there is a placeholder. A prompt injection, a
log line, an exception dump, or a /v1/model/info response can only ever
contain that.
Three details to get right, each of which fails confusingly on its own:
- Forward mode, not reverse. One LiteLLM process egresses to every provider
in its
model_list, so the credential has to be swapped on the way out rather than at a base URL the gateway was pointed at. (Reverse mode does work — drop--mode forwardand you get a route per provider — but then every model entry needs its ownapi_base.) SSL_CERT_FILE, notNODE_EXTRA_CA_CERTS. The generated config's env hints name the Node variable, because it cannot know the runtime. LiteLLM is Python: httpx and requests readSSL_CERT_FILEandREQUESTS_CA_BUNDLE, and neither reads the Node one. Getting it wrong surfaces as a certificate verification failure against the provider, which reads like a proxy bug.- Each key is fenced to its own host. The generated rules permit
OPENAI_API_KEYtowardapi.openai.comand nowhere else. A gateway holding three provider keys cannot send one upstream that should never see it, and a provider added toconfig.yamlbut not to the proxy fails closed.
Do not resolve a provider key through the secret manager as well. A key the
manager resolves is a key in the gateway's memory, which is exactly what this
section removes. The two combine the other way round: the proxy for provider
keys, the manager for the gateway's own DATABASE_URL and
LITELLM_MASTER_KEY.
Which one
| If… | Use |
|---|---|
You just don't want a .env file next to the gateway | Wrap the process |
| The platform gives you no way to wrap the process | The secret manager |
| A rotated key should reach a running gateway without a restart | The secret manager |
| The gateway is multi-tenant, or runs code you didn't write | The proxy |
| You want the gateway's own credentials in seekrit too | The secret manager, alongside the proxy |
All three read the same environment through the same kind of token, so moving between them is a deployment change and never a key rotation.