Providers & Auth
rupu speaks to four built-in model providers at once — plus any OpenAI-compatible endpoint you register — and because every agent picks its own provider, model, and credential, a single workflow can run a planner on one model and reviewers on another in the same run.
The multi-provider model
rupu is not pinned to one vendor. It ships working integrations for Anthropic, OpenAI, Gemini, and GitHub Copilot, and the choice is made per agent, not globally. Each agent file declares its provider, model, and (optionally) auth mode in its YAML frontmatter — so when a workflow fans out across several agents, those agents can each talk to a different provider concurrently.
That means a real workflow might dispatch a planner step on a Claude model, two reviewer steps on GPT-5 and Gemini, and a long-context search step on Gemini's 1M-token window — all inside one run, each authenticating with its own credential.
| Provider | What it's for | Auth options |
|---|---|---|
anthropic |
Claude models. The most-exercised provider in rupu. | Console API key or Claude.ai SSO (browser callback). |
openai |
GPT models via the Responses API. | Platform API key or ChatGPT SSO (browser callback). Different endpoints under the hood. |
gemini |
Gemini models with very long context windows. | SSO via Vertex / Gemini CLI OAuth (browser callback). API key via AI Studio is deferred. |
copilot |
Copilot-hosted models. Requires a paid Copilot subscription. | GitHub PAT (GITHUB_TOKEN) or GitHub device-code SSO. |
Those four are the built-in providers, with first-class auth flows and curated model lists. But you are not limited to them: beyond the built-ins, you can register any OpenAI-compatible HTTP endpoint as a named provider and mix it into the same workflows — see Bring your own provider below.
Choosing a provider/model per agent
Provider and model are agent-level decisions. An agent file is Markdown with a YAML frontmatter block; the relevant fields are provider, model, and the optional auth:
--- name: refactor description: Suggest minimal-diff refactors. provider: anthropic auth: sso model: claude-sonnet-4-6 --- You suggest minimal-diff refactors. Be concise.
If you omit model, rupu falls back to the provider's default_model from ~/.rupu/config.toml. If you omit auth, the credential resolver applies a default precedence (SSO if present and refreshable, then API key).
Model resolution runs most-specific-wins: the agent's (or step's) model: first, then the resolved provider's own [providers.<name>].default_model, then the global top-level default_model in config.toml, then a baked-in fallback. So a provider-scoped default_model beats the global one — an agent pinned to a custom endpoint with no model: of its own gets that provider's model, not the global default (which the custom endpoint would typically reject as unknown).
Authentication
Under the hood rupu has two neutral auth modes — api-key and sso — and the SSO mode resolves to the right OAuth flow per provider (a localhost browser callback for Anthropic / OpenAI / Gemini, or a GitHub device code for Copilot).
| Mode | CLI flag | How it works | Providers |
|---|---|---|---|
| API key | --mode api-key |
Paste the secret with --key or pipe it on stdin. Stored under the key <provider>/api-key in ~/.rupu/auth.json. |
all four |
| SSO — browser callback | --mode sso |
Binds a localhost listener, opens the provider's authorize URL with a PKCE challenge, validates state, exchanges the code for tokens. No headless fallback. |
anthropic, openai, gemini |
| SSO — device code | --mode sso |
Prints a code + URL; you authorize from any browser anywhere while rupu polls. Works headless / over SSH. | copilot |
SSO access tokens (typically ~1 hour) are refreshed pre-emptively — the resolver refreshes when expires_at - now < 60s on a credential read, using the stored refresh token. There is no automatic fall-back from SSO to API key: if you chose SSO, a refresh failure points you back at rupu auth login.
How credentials are stored
There is one credential backend: a JSON file. The OS keychain backend was retired — a bare CLI binary's keychain access requirement is cdhash-bound, so every rebuild invalidated it and the next read failed silently ("my credentials vanished after an update"). Peer CLIs — gh, aws, gcloud, kubectl, terraform — store credentials in files for the same reason.
- The store —
~/.rupu/auth.json, mode 0600. A flat JSON object of<provider>/<mode>keys (e.g.anthropic/api-key,gemini/sso) to secret payloads. Permissions are re-applied to0600on every write, so a previously loose file is tightened the next time you log in; a wider mode on read logs a warning rather than refusing — refusing would block recovery on a misconfigured machine. - Plaintext, so the file mode is the protection. Secrets are not encrypted at rest and are never printed back by
rupu auth status. Keep~/.rupuout of backups you would not trust with an API key. - Relocatable —
RUPU_AUTH_FILEandRUPU_HOME.RUPU_AUTH_FILEpoints at an exact file;RUPU_HOMEmoves the whole rupu directory ($RUPU_HOME/auth.json). These select a path — nothing selects a backend any more, and the oldRUPU_AUTH_BACKENDvariable no longer diverts storage. - Global only.
auth.jsonis never read from a project directory — credentials live at the user level.
rupu auth login probes for them (read-only, it never reads a secret or prompts for access) and prints a notice naming the affected providers — built-ins and your own openai-compatible ones alike. The remedy is to re-authenticate: rupu auth login --provider <name> --mode api-key|sso, which writes to the file store. Nothing is imported or deleted automatically; the old entries stay untouched and you can clear each one with security delete-generic-password -s rupu -a <account>.The rupu auth CLI
The rupu auth subcommand manages stored credentials:
rupu auth login— store credentials for a provider (--provider,--mode api-key|sso, optional--key). Also where the stranded-keychain notice fires.rupu auth logout— remove a stored credential (--provider[--mode], or--all[--yes]).rupu auth status— show which providers have an API key or SSO credential stored; never prints secrets.rupu auth backend— report where credentials live. It no longer switches anything:--use fileis the only accepted value (and is already in effect), and--use keychainfails with an explanation rather than silently writing to a file you did not choose.
# API key: pass it inline, or pipe it on stdin rupu auth login --provider anthropic --mode api-key --key sk-ant-XXX echo -n "$KEY" | rupu auth login --provider openai --mode api-key # SSO (browser callback or device code, chosen per provider) rupu auth login --provider copilot --mode sso # Inspect and clean up rupu auth status rupu auth logout --provider gemini --mode sso rupu auth logout --all --yes # Report where credentials live (there is only one backend) rupu auth backend # Store the file somewhere else export RUPU_AUTH_FILE=/secure/volume/rupu-auth.json
Models are discovered separately with rupu models: rupu models list shows the catalog across all four providers (with a --provider filter), and rupu models refresh re-fetches the live caches. An agent's model: value is resolved against custom config entries, the live cache, then a baked-in list.
Per-provider notes
Anthropic
Console API keys start with sk-ant- and are shown only once — copy at creation. SSO authenticates via Claude.ai's OAuth (claude.ai/oauth/authorize), good for Claude Pro subscribers. Access tokens last ~1 hour and refresh automatically.
OpenAI
Two paths hit different endpoints: a Platform API key (sk-...) targets api.openai.com/v1/responses and is billed via OpenAI; ChatGPT SSO targets the ChatGPT backend and is covered by a Plus/Pro subscription. rupu detects the chatgpt_account_id claim and routes accordingly. Org-scoped keys need org_id in config. Only the newer Responses API is supported — pre-Responses models won't work.
Gemini
Currently SSO only: the lifted client targets the Vertex AI / Gemini CLI OAuth path (accounts.google.com, cloud-platform scope). API-key auth via AI Studio is deferred (--mode api-key returns NotWiredInV0). You need a Google Cloud project with the Vertex AI API enabled and billing configured; set region in config for Vertex.
Copilot
Requires a paid Copilot Pro / Business / Enterprise subscription — a free GitHub account can't invoke the API even with a valid token. The simplest path is a GitHub PAT (ghp_...), or pick it up from GITHUB_TOKEN when --key is omitted. The device-code SSO flow is the same one gh auth login uses and is the only SSO flow that works headless. The exchanged Copilot token expires faster than the GitHub PAT; rupu re-mints it internally.
Bring your own provider (OpenAI-compatible)
The four built-ins are not the boundary. With the openai-compatible provider kind you can register any HTTP server that speaks the OpenAI /v1/chat/completions API as a named provider, then use it from agents exactly like a built-in — and mix it with the built-ins in one workflow. That covers:
- Local models — vLLM, llama.cpp, or Ollama running on your own machine or a GPU box.
- Gateways & aggregators — OpenRouter, Together, Groq, Fireworks, and similar OpenAI-shaped endpoints.
- Private / enterprise endpoints — Oracle GenAI or an internal company gateway.
Authentication is a single static Bearer API key per provider — there is no SSO flow for openai-compatible providers. The model catalog is whatever you declare in config; rupu does not call /v1/models on these endpoints.
Configure it
Add a [providers.<name>] table to ~/.rupu/config.toml with kind = "openai-compatible". base_url and default_model are required; stream is optional (defaults to true; set false for servers without an SSE endpoint). Add one [[providers.<name>.models]] block per private or fine-tuned model you want to select.
default_provider = "oracle" [providers.oracle] kind = "openai-compatible" base_url = "http://192.29.35.246:8080" default_model = "/raid/models/zai-org/GLM-5.2-FP8" stream = true # set false if the server has no SSE endpoint [[providers.oracle.models]] id = "/raid/models/zai-org/GLM-5.2-FP8" context_window = 131072 max_output = 8192
base_url may include or omit a trailing /v1 — rupu normalises both to <root>/v1/chat/completions. Each [[providers.<name>.models]] entry requires id; context_window and max_output are optional (defaulting to 32768 and 8192). Name the provider anything you like — oracle, vllm, together — except a reserved built-in name (anthropic, openai, gemini, copilot); rupu rejects the config otherwise.
Authenticate
Supply the static Bearer key with rupu auth login (stored in ~/.rupu/auth.json under <name>/api-key, same as the built-ins), or — for CI and ephemeral environments — set the env var RUPU_<UPPERCASED_PROVIDER_NAME>_API_KEY and rupu reads it automatically.
# Interactive prompt (the key is not echoed), or pipe it on stdin rupu auth login --provider oracle --mode api-key # Or set the env var directly — pattern: RUPU_<PROVIDER>_API_KEY export RUPU_ORACLE_API_KEY=sk-...
Use it from an agent
Point an agent at the provider by name; the model: can be any id the endpoint serves, including the custom [[providers.<name>.models]] entries that make private or fine-tuned models selectable.
--- name: oracle-codereview description: Code review via Oracle GenAI. provider: oracle model: /raid/models/zai-org/GLM-5.2-FP8 --- You review code changes for correctness, style, and missing tests.
Then run it like any other agent, and confirm the custom catalog with rupu models list --provider oracle (entries show source custom):
rupu run --agent oracle-codereview
Example endpoints
Any of these can be dropped into the base_url above. When a path detail is unknown for your deployment, the generic http://host:port/v1 shape is a safe starting point.
| Target | Rough base_url |
|---|---|
| Ollama (local) | http://localhost:11434/v1 |
| vLLM (self-hosted) | http://host:8000/v1 |
| llama.cpp server | http://host:8080/v1 |
| OpenRouter | https://openrouter.ai/api/v1 |
| Together | https://api.together.xyz/v1 |
| Groq | https://api.groq.com/openai/v1 |
| Fireworks | https://api.fireworks.ai/inference/v1 |
| Oracle GenAI / internal gateway | http://host:port |
$0.00 in cost tracking (no pricing tables), and model listing returns only what you declare under [[providers.<name>.models]] — rupu never queries /v1/models on these endpoints.