* feat(provider): add Cloudflare Workers AI integration
Adds Cloudflare Workers AI as a first-class OpenAI-compatible provider
preset, modeled on the Venice / Xiaomi MiMo descriptors.
- New `src/integrations/vendors/cloudflare.ts` descriptor:
- `classification: 'openai-compatible'`
- Default base URL with literal `<ACCOUNT_ID>` placeholder — users
substitute via `/provider` baseUrl edit, same shape as the Azure
OpenAI example already in `docs/advanced-setup.md`
- `CLOUDFLARE_API_TOKEN` env, with `OPENAI_API_KEY` as fallback
- `removeBodyFields: ['store']` since Workers AI rejects unknown
OpenAI body fields (mirrors Mistral / Gemini / Cerebras strip)
- Static catalog with current Workers AI chat models
(`@cf/meta/llama-3.3-70b-instruct-fp8-fast`,
`@cf/meta/llama-3.1-8b-instruct`,
`@cf/deepseek-ai/deepseek-r1-distill-qwen-32b`,
`@cf/qwen/qwen2.5-coder-32b-instruct`)
- Validation routing on `api.cloudflare.com` /
`gateway.ai.cloudflare.com` hosts so an env-pasted URL maps back
to the preset
- Env mirror sites in `src/utils/providerProfiles.ts`: mirror api key
into `CLOUDFLARE_API_TOKEN` when baseUrl contains a Cloudflare host
(3 sites: same-env check, openAIProfileEnv build, applyEnv).
- `CLOUDFLARE_API_TOKEN` added to `PROFILE_ENV_KEYS` / `SECRET_ENV_KEYS` /
`ProfileEnv` / `SecretValueSource` in `src/utils/providerProfile.ts`
so the profile-clean and secret-redact paths know about it.
- `src/utils/providerFlag.ts` `--provider <name>` startup flag now
detects a Cloudflare profile from `OPENAI_API_KEY ===
CLOUDFLARE_API_TOKEN` (mirrors how the other host-key mirrors are
reverse-mapped to their preset id).
- `bun run scripts/generate-integrations-artifacts.ts` regenerated
`integrationArtifacts.generated.ts` to include the cloudflare preset
+ route + vendor.
- Tests: `compatibility.test.ts` PRESETS list, new
`buildProfileSaveMessage` Cloudflare case in `provider.test.tsx`,
new `applyProviderProfileToProcessEnv` Cloudflare case in
`providerProfiles.test.ts`.
- Docs: README providers table row + `docs/advanced-setup.md` section
matching the MiMo / Mistral entries.
- Dedicated AI Gateway integration with `gateway_id` URL templating.
Today users can still paste a full Gateway URL into `OPENAI_BASE_URL`
and the preset's `matchBaseUrlHosts` picks `gateway.ai.cloudflare.com`
up.
- Dynamic `/models` discovery on the Groq #1143 / `mapModel` pattern —
Cloudflare's `/v1/models` returns the runnable model list and the
hybrid catalog path drops in cleanly. Left as a separate PR so this
one stays a focused preset add.
Closes #1100
* fix(cloudflare): narrow route matching to api.cloudflare.com host
`gateway.ai.cloudflare.com` is the shared host for *all* Cloudflare AI
Gateway routes (Workers AI, Anthropic, OpenAI, etc.), so matching it to
the Workers AI preset applied Workers-AI runtime metadata and
credential precedence (CLOUDFLARE_API_TOKEN before OPENAI_API_KEY, body
'store' strip, max_tokens field) to other providers' Gateway URLs.
Drop the shared host from the match list; a dedicated AI Gateway
integration with path-aware routing is the right follow-up.
Refs #1100.
* fix(provider-manager): keep Codex OAuth after DeepSeek when cloudflare added
The picker hardcoded `options.splice(7, 0, …)` to drop the Codex OAuth
entry right after DeepSeek. Adding cloudflare to ORDERED_PROVIDER_PRESETS
bumped DeepSeek to index 7, so the splice now lands Codex OAuth *before*
DeepSeek and breaks the test fixture that drives navigateToPreset by
keypress count.
Switch to a dynamic `findIndex('deepseek') + 1` lookup so any future
preset inserted between Bankr and DeepSeek keeps the established
ordering. Fixture updated to mirror the new picker order.
Caught by CI on 12b3ff… smoke-and-tests: 8 ProviderManager tests
timing out because navigateToPreset overshot/undershot the target.
* fix(cloudflare): exclude the shared AI Gateway host from Cloudflare routing
The profile env/alignment/startup paths mirrored CLOUDFLARE_API_TOKEN whenever
the profile URL merely contained 'gateway.ai.cloudflare.com'. That host is the
shared AI Gateway for all Cloudflare AI routes (Workers AI, OpenAI, Anthropic,
...), so a profile retargeted to /openai or /anthropic Gateway URLs was wrongly
tied to the Cloudflare route and credential precedence.
Add isCloudflareBaseUrl (hostname === api.cloudflare.com, matching the Workers
AI host and the descriptor's matchBaseUrlHosts) and route all three sites
through it, consistent with isXaiBaseUrl/isFireworksBaseUrl. Also restore
CLOUDFLARE_API_TOKEN in the provider profile test cleanup keys.
* fix(cloudflare): don't seed the placeholder base URL from the CLI shortcut
`openclaude --provider cloudflare` fell through the generic OpenAI-compatible
branch and applied the descriptor default base URL verbatim — including the
unresolved `<ACCOUNT_ID>` placeholder — leaving the shortcut 'configured' with
an endpoint that cannot serve a request. Skip seeding any base URL that still
contains a `<...>` placeholder, so the user must supply a real account-scoped
URL (OPENAI_BASE_URL / `/provider` edit) first, matching how the wizard treats
placeholder endpoints.
* test(cloudflare): assert exact null fallback for AI Gateway routes
The shared AI Gateway URL assertions used `.not.toBe('cloudflare')`, which
would also pass for any other non-cloudflare return value. The intended
fallback is null, so assert `.toBe(null)` to lock the regression boundary.
* fix(cloudflare): gate profile token mirroring on base URL host only
applyProviderProfileToProcessEnv mirrored CLOUDFLARE_API_TOKEN whenever
route.routeId === 'cloudflare'. route comes from the saved profile.provider,
so that disjunct is always true for a cloudflare profile, including one
retargeted to the shared gateway.ai.cloudflare.com AI Gateway host. The
sibling sites (isProcessEnvAlignedWithProfile, buildOpenAICompatibleStartupEnv)
already key on isCloudflareBaseUrl only; align this site with them so a
shared-gateway profile no longer leaks the token or stays pinned to the
cloudflare route. Add a regression test for the gateway.ai.cloudflare.com case.
* chore(integrations): regenerate artifacts for the Cloudflare vendor
The rebase took main's generated artifacts at the conflict; regenerate so the
Cloudflare vendor descriptor is registered in VENDOR_DESCRIPTORS and the
manifest alongside the providers main added.
* fix(cloudflare): mirror CLOUDFLARE_API_TOKEN into the OpenAI-compatible auth path
The --provider cloudflare shortcut fell through to the generic
OpenAI-compatible default branch and never copied CLOUDFLARE_API_TOKEN
into OPENAI_API_KEY, so a user who only set the token sent an
unauthenticated request. Add a dedicated cloudflare case that mirrors the
token (and clears a stale generic key when absent), keeping the
placeholder-URL skip.
buildOpenAICompatibleStartupEnv also returned from its strict-env branch
before the fallback CLOUDFLARE_API_TOKEN mirror, so a keyed Cloudflare
profile persisted a startup env that omitted the token and re-detected
inconsistently after relaunch. Mirror it in the strict branch alongside
nearai/fireworks. Add regression coverage for both paths.
* fix(cloudflare): gate token mirroring on a real Cloudflare endpoint
The cloudflare shortcut copied CLOUDFLARE_API_TOKEN into the generic
OPENAI_API_KEY unconditionally. The descriptor default carries an
unresolved `<ACCOUNT_ID>` placeholder and is never seeded, so with
OPENAI_BASE_URL unset (or still pointing at a previous OpenAI-compatible
provider) the token would be attached to the wrong host. Gate the mirror
on isCloudflareBaseUrl(getConfiguredOpenAIBaseUrl()) — only seed
OPENAI_API_KEY once the configured base URL resolves to
api.cloudflare.com, otherwise fail fast and leave it unset. Add
regression coverage for the unconfigured, stale-host, and AI-Gateway-host
cases.
* fix(cloudflare): reject placeholder URL and keep the OPENAI_API_KEY fallback
The token mirror keyed on the api.cloudflare.com host only, so the literal
<ACCOUNT_ID> placeholder URL (same host) passed the gate and copied the
token onto a non-working endpoint. It also deleted any generic
OPENAI_API_KEY when no token was set, breaking the documented
compatibility fallback for users authenticating a real Workers AI URL with
OPENAI_API_KEY. Mirror only on a real (non-placeholder) Cloudflare
endpoint, and preserve an existing generic key there when no dedicated
token is present.
Refs #1100
* refactor(cloudflare): model Workers AI as a gateway, not a vendor
Cloudflare Workers AI is a hosted OpenAI-compatible inference endpoint
reached over the shared openai transport, so it belongs with the gateway
providers (atlas-cloud, groq, together, ...) rather than the transport
vendors. Move it to gateways/cloudflare.ts via defineGateway (category
hosted, vendorId openai), regenerate the integration artifacts, and
allowlist its provider-specific @cf/* catalog ids in the gateway
descriptor check (no shared cross-provider descriptor exists, same as
azure-deployment).
Refs #1100
* fix(cloudflare): key Workers AI detection on the account path, not the host
api.cloudflare.com also serves the general Cloudflare REST API, so matching the
whole host treated unrelated URLs (e.g. /client/v4/user/tokens/verify) as the
Workers AI route and mirrored CLOUDFLARE_API_TOKEN into OPENAI_API_KEY for them.
isCloudflareBaseUrl now requires the Workers AI path
/client/v4/accounts/<account_id>/ai/v1 with a real (non-placeholder) account id,
and resolveRouteIdFromBaseUrl guards its cloudflare hostname match through the
same predicate. Both route detection and token/profile mirroring key on the
actual Workers AI endpoint.
Adds same-host negative regressions (general REST path is not routed and does
not mirror the token; unresolved <ACCOUNT_ID> placeholder is excluded) and
asserts the Cloudflare Workers AI preset appears in the first-run picker.
* fix(cloudflare): honor the Workers AI path boundary in the profile-provider fallback
resolveActiveRouteIdFromEnv returned the saved active-profile provider's route
id before consulting its base URL. For a `cloudflare` profile that had been
retargeted to a non-Workers URL — the shared AI Gateway host, or a general
api.cloudflare.com REST path — this still resolved as `cloudflare`, so the
Workers AI shim config (removeBodyFields: ['store'], Cloudflare model metadata)
and CLOUDFLARE_API_TOKEN mirroring were applied to a generic endpoint, even
though resolveRouteIdFromBaseUrl already excludes those URLs.
Gate the profile-provider shortcut through profileRouteHonorsBaseUrlBoundary,
which requires the path-aware isCloudflareBaseUrl for the cloudflare route (all
other routes are host-scoped by resolveProfileRoute and unaffected). A retargeted
profile now falls through to the generic openai/custom resolution; a genuine
Workers AI profile base URL still resolves as cloudflare.
Adds regressions for both retarget cases (gateway host + REST path) and the
positive Workers AI profile case.
* fix(cloudflare): require HTTPS and honor the Workers AI path in validation
isCloudflareBaseUrl accepted any scheme, so http://api.cloudflare.com/
client/v4/accounts/<id>/ai/v1 resolved as the cloudflare route and mirrored
CLOUDFLARE_API_TOKEN into OPENAI_API_KEY over cleartext. Require url.protocol
=== 'https:'.
Startup validation selected the Cloudflare target on host match alone, so a
non-Workers path like /client/v4/user/tokens/verify demanded Workers AI auth
instead of falling back to generic OpenAI validation. Gate the cloudflare
target on isCloudflareBaseUrl(request.baseUrl), mirroring the runtime route
resolver's path boundary.
* test(cloudflare): lock non-Workers path token boundary; fix stale host-only comments
The apply/persist paths already gate CLOUDFLARE_API_TOKEN mirroring on the
isCloudflareBaseUrl path predicate, but had no coverage for a same-host
non-Workers path (api.cloudflare.com/client/v4/user/tokens/verify) and the
comments beside the mirroring sites still described a host-only boundary.
Add negative apply and persist regressions asserting the token is not mirrored
or persisted for that non-Workers URL, and update the comments to describe the
real Workers AI path predicate instead of host-only matching.
* fix(cloudflare): fall back to a generic route for retargeted profiles
resolveProfileCapabilityRouteId returned the cloudflare capability route id for
any saved cloudflare profile whose base URL no longer resolves — including one
retargeted to gateway.ai.cloudflare.com or another OpenAI-compatible host. That
stripped generic capabilities (apiFormat, custom auth/request headers) from
profile sanitize/apply even though the runtime resolver runs such a profile as
a generic OpenAI-compatible route. Mirror the same isCloudflareBaseUrl boundary:
keep the cloudflare route only for the real Workers AI URL (or the unset
descriptor default) and fall back to 'custom' otherwise. Regression asserts a
retargeted cloudflare profile preserves OPENAI_API_FORMAT.
* test(cloudflare): assert retargeted profile resolves to the custom route
Pin both resolveActiveRouteIdFromEnv assertions for a retargeted cloudflare
profile to .toBe('custom') instead of .not.toBe('cloudflare'), so the test
locks the intended generic OpenAI-compatible fallback rather than merely
excluding the cloudflare route.
29 KiB
OpenClaude Advanced Setup
This guide is for users who want source builds, Bun workflows, provider profiles, diagnostics, or more control over runtime behavior.
Install Options
OpenClaude requires Node.js >=22.0.0 for npm installs and runtime. Bun is
only required when building or running from source.
Option A: npm
npm install -g @gitlawb/openclaude@latest
Option B: From source with Bun
Use Bun 1.3.13 or newer for source builds. Older Bun versions can fail during bun run build.
git clone https://github.com/Gitlawb/openclaude.git
cd openclaude
bun install
bun run build
npm link
Option C: Run directly with Bun
git clone https://github.com/Gitlawb/openclaude.git
cd openclaude
bun install
bun run dev
Provider Examples
OpenAI
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=sk-...
export OPENAI_MODEL=gpt-4o
Codex via ChatGPT auth
codexplan maps to GPT-5.5 on the Codex backend with high reasoning.
codexspark maps to GPT-5.3 Codex Spark for faster loops.
If you use the in-app provider wizard, choose Codex OAuth to open ChatGPT sign-in in your browser and let OpenClaude store Codex credentials securely.
If you already use the Codex CLI, OpenClaude reads ~/.codex/auth.json automatically. You can also point it elsewhere with CODEX_AUTH_JSON_PATH or override the token directly with CODEX_API_KEY.
If you set CODEX_API_KEY manually and are not relying on auth.json or stored
Codex OAuth credentials, also set CHATGPT_ACCOUNT_ID (or
CODEX_ACCOUNT_ID).
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_MODEL=codexplan
# optional if you do not already have ~/.codex/auth.json
export CODEX_API_KEY=...
export CHATGPT_ACCOUNT_ID=...
openclaude
DeepSeek
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=sk-...
export OPENAI_BASE_URL=https://api.deepseek.com/v1
export OPENAI_MODEL=deepseek-v4-flash
Use deepseek-v4-pro when you want the stronger model. deepseek-chat and deepseek-reasoner remain available as DeepSeek's legacy API aliases.
Google Gemini
export CLAUDE_CODE_USE_GEMINI=1
export GEMINI_API_KEY=...
export GEMINI_MODEL=gemini-3-flash-preview
Claude on Vertex AI
The Vertex route uses Anthropic's Claude-on-Vertex API. It is not a general Vertex AI Model Garden adapter for Gemini or arbitrary partner models; use the Gemini provider for Gemini models and OpenAI-compatible routes for compatible third-party gateways.
Authentication uses Google Application Default Credentials through
google-auth-library. There is no OPENAI_API_KEY-style API key for this
route. For global npm installs, install the auth package on demand (it is
not bundled by default — see Optional provider packages):
npm i -g google-auth-library
Authenticate with either local Application Default Credentials (ADC) or a service-account key file:
# Option 1 — local ADC (interactive, uses your own Google account):
gcloud auth application-default login
# Option 2 — service-account key file (headless / CI):
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account.json
Minimal setup:
export CLAUDE_CODE_USE_VERTEX=1
export ANTHROPIC_VERTEX_PROJECT_ID=my-gcp-project
export GOOGLE_CLOUD_PROJECT=my-gcp-project
export CLOUD_ML_REGION=us-east5
openclaude --model claude-sonnet-4-6
CLOUD_ML_REGION is optional and defaults to us-east5. Model-specific
Vertex region override variables are also supported for Claude models; see
src/utils/envUtils.ts for the current override names.
Gemini via OpenRouter
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=sk-or-...
export OPENAI_BASE_URL=https://openrouter.ai/api/v1
export OPENAI_MODEL=google/gemini-2.5-pro
OpenRouter model availability changes over time. If a model stops working, try another current OpenRouter model before assuming the integration is broken.
Ollama
ollama pull llama3.3:70b
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_BASE_URL=http://localhost:11434/v1
export OPENAI_MODEL=llama3.3:70b
Ollama Context Length
OpenClaude sends the current conversation history to Ollama on each turn and
uses Ollama's native chat API for Ollama endpoints. Native chat lets OpenClaude
send options.num_ctx with each request, so Ollama receives a 32768-token
context window by default instead of falling back to the smaller context often
used by Ollama's OpenAI-compatible /v1/chat/completions shim.
To choose a different request-level context size, set
OPENCLAUDE_OLLAMA_NUM_CTX before launching OpenClaude:
export OPENCLAUDE_OLLAMA_NUM_CTX=65536
You can also start Ollama with a global context length:
macOS / Linux:
# Stop any existing Ollama app/server first, then run:
OLLAMA_CONTEXT_LENGTH=32768 ollama serve
Windows PowerShell:
# Quit any existing Ollama app/server first, then run:
$env:OLLAMA_CONTEXT_LENGTH="32768"
ollama serve
After a chat request, verify the loaded model is using the requested context:
ollama ps
Check the CONTEXT column. If it still shows a small value such as 4K after a
new OpenClaude request, stop the existing Ollama app/server, start it again, and
retry the request.
Use a concrete recall test after changing the setting, such as asking the model to repeat the first topic from the current chat. Questions like "do you remember our conversation?" can trigger generic local-model disclaimers even when history is present.
Atomic Chat (local, Apple Silicon)
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_BASE_URL=http://127.0.0.1:1337/v1
export OPENAI_MODEL=your-model-name
No API key is needed for Atomic Chat local models.
Or use the profile launcher:
bun run dev:atomic-chat
Download Atomic Chat from atomic.chat. The app must be running with a model loaded before launching.
LM Studio
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_BASE_URL=http://localhost:1234/v1
export OPENAI_MODEL=your-model-name
Together AI
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=...
export OPENAI_BASE_URL=https://api.together.xyz/v1
export OPENAI_MODEL=meta-llama/Llama-3.3-70B-Instruct-Turbo
Groq
export CLAUDE_CODE_USE_OPENAI=1
export GROQ_API_KEY=gsk_...
export OPENAI_BASE_URL=https://api.groq.com/openai/v1
export OPENAI_MODEL=llama-3.3-70b-versatile
GROQ_API_KEY matches the built-in Groq gateway preset. OPENAI_API_KEY also works as a fallback on the generic OpenAI-compatible path, but GROQ_API_KEY is the preferred variable for Groq-specific setup.
OpenCode Zen (pay-as-you-go)
export CLAUDE_CODE_USE_OPENAI=1
export OPENCODE_API_KEY=...
export OPENAI_BASE_URL=https://opencode.ai/zen/v1
export OPENAI_MODEL=gpt-5.4
openclaude
OpenCode Zen is a pay-as-you-go AI gateway with 48 models (GPT, Claude, Gemini,
Qwen, MiniMax, GLM, Kimi, Grok, Big Pickle, DeepSeek, Nemotron). Uses the same
OPENCODE_API_KEY as OpenCode Go. Get your key from https://opencode.ai.
OpenCode Go (subscription)
export CLAUDE_CODE_USE_OPENAI=1
export OPENCODE_API_KEY=...
export OPENAI_BASE_URL=https://opencode.ai/zen/go/v1
export OPENAI_MODEL=glm-5.1
openclaude
OpenCode Go is a $10/mo subscription for 13 open models (GLM, Kimi, DeepSeek,
MiMo, MiniMax, Qwen). Uses the same OPENCODE_API_KEY as OpenCode Zen.
Gitlawb Opengateway
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_BASE_URL=https://opengateway.gitlawb.com/v1
export OPENGATEWAY_API_KEY=ogw_live_...
export OPENAI_MODEL=mimo-v2.5-pro
The Opengateway route is the fresh-install startup default and requires an API
key from https://gitlawb.com/opengateway/keys. Keep the base URL at /v1 and
switch models with /model or OPENAI_MODEL. Current partner models include:
mimo-v2.5-progoogle/gemini-3.1-flash-lite-preview
Xiaomi MiMo
export CLAUDE_CODE_USE_OPENAI=1
export MIMO_API_KEY=...
export OPENAI_BASE_URL=https://api.xiaomimimo.com/v1
export OPENAI_MODEL=mimo-v2.5-pro
The /provider Xiaomi MiMo preset uses the same endpoint and stores the key as MIMO_API_KEY. OPENAI_API_KEY also works as a compatibility fallback, but MIMO_API_KEY keeps the profile tied to the MiMo route.
NEAR AI
export CLAUDE_CODE_USE_OPENAI=1
export NEARAI_API_KEY=...
export OPENAI_BASE_URL=https://cloud-api.near.ai/v1
export OPENAI_MODEL=anthropic/claude-sonnet-4-6
openclaude
NEAR AI is a unified OpenAI-compatible gateway that proxies Anthropic, OpenAI, and Google models alongside TEE-hosted open models (GLM 5.1, Qwen3.5, Kimi K2.6). All models are accessible from a single endpoint with one API key. Get your key from https://cloud.near.ai/dashboard/organizations.
Model IDs use provider/model-name format (e.g. anthropic/claude-opus-4-7,
openai/gpt-5.5, google/gemini-3.5-flash, zai-org/GLM-5.1-FP8).
For direct TEE completions (lower latency, verifiable privacy):
export OPENAI_BASE_URL=https://qwen35-122b.completions.near.ai/v1
Cloudflare Workers AI
export CLAUDE_CODE_USE_OPENAI=1
export CLOUDFLARE_API_TOKEN=...
export OPENAI_BASE_URL=https://api.cloudflare.com/client/v4/accounts/<ACCOUNT_ID>/ai/v1
export OPENAI_MODEL=@cf/meta/llama-3.3-70b-instruct-fp8-fast
Replace <ACCOUNT_ID> with your Cloudflare account id (visible in the Cloudflare dashboard URL). OPENAI_API_KEY also works as a compatibility fallback, but CLOUDFLARE_API_TOKEN keeps the profile tied to the Cloudflare preset. The /provider Cloudflare Workers AI preset stores the token under CLOUDFLARE_API_TOKEN.
Mistral
export CLAUDE_CODE_USE_MISTRAL=1
export MISTRAL_API_KEY=...
export MISTRAL_MODEL=devstral-latest
Azure OpenAI
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=your-azure-key
export OPENAI_BASE_URL=https://your-resource.openai.azure.com/openai/deployments/your-deployment/v1
export OPENAI_MODEL=gpt-4o
Microsoft Foundry / Azure OpenAI (resource URL + deployment)
When your endpoint is the resource base URL (not the full .../deployments/.../v1 path), set OPENAI_MODEL to the deployment name and AZURE_OPENAI_API_VERSION to your API version. The OpenAI shim builds:
{base}/openai/deployments/{OPENAI_MODEL}/chat/completions?api-version={AZURE_OPENAI_API_VERSION}
and sends the key in the api-key header for Azure hosts.
export CLAUDE_CODE_USE_OPENAI=1
export OPENAI_API_KEY=your-azure-key
export OPENAI_BASE_URL=https://your-resource.openai.azure.com
export OPENAI_MODEL=your-deployment-name
export AZURE_OPENAI_API_VERSION=2024-12-01-preview
If your hostname is not detected as Azure (for example some inference endpoints), force Azure URL and header behavior:
export OPENAI_AZURE_STYLE=1
Fireworks AI
Fireworks AI provides a fully OpenAI-compatible endpoint. Model IDs use the full path format accounts/fireworks/models/<model-name>.
export CLAUDE_CODE_USE_OPENAI=1
export FIREWORKS_API_KEY=fw_your_key_here
export OPENAI_BASE_URL=https://api.fireworks.ai/inference/v1
export OPENAI_MODEL=accounts/fireworks/models/llama-v3p1-70b-instruct
The OpenClaude VS Code extension can store the key in Secret Storage and set these variables for you when you launch from the Control Center. See vscode-extension/openclaude-vscode/README.md.
Optional provider packages
To keep the default npm i -g @gitlawb/openclaude install small and
warning-free, a few provider SDKs and the native image library are not
bundled. They are loaded on demand, and the CLI prints an npm install <pkg>
hint (add -g for the global CLI) if you enable a feature whose package is
missing. Install only what you need:
| Feature | Trigger | Install |
|---|---|---|
| AWS Bedrock | CLAUDE_CODE_USE_BEDROCK=1 |
npm i -g @anthropic-ai/bedrock-sdk. Profile-based auth (~/.aws/credentials) additionally needs @aws-sdk/credential-providers and @aws-sdk/client-sts; model listing needs @aws-sdk/client-bedrock. Proxy and skip-auth setups may also need @aws-sdk/credential-provider-node, @smithy/node-http-handler, or @smithy/core. The CLI prints the exact missing package if you hit one. |
| Azure Foundry | CLAUDE_CODE_USE_FOUNDRY=1 |
npm i -g @anthropic-ai/foundry-sdk @azure/identity |
| Claude on Vertex AI / Gemini ADC | CLAUDE_CODE_USE_VERTEX=1 / Gemini ADC auth |
npm i -g google-auth-library |
| Reading/processing images | reading an image file | npm i -g sharp |
When installing OpenClaude from source (bun install), all of these are
already present as dev dependencies, so source/dev builds need no extra steps.
Environment Variables
| Variable | Required | Description |
|---|---|---|
CLAUDE_CODE_USE_OPENAI |
OpenAI-compatible only | Set to 1 to enable the OpenAI-compatible provider path |
OPENAI_API_KEYS |
One of OPENAI_API_KEYS or OPENAI_API_KEY for non-local OpenAI-compatible cloud routes* |
Comma-separated OpenAI-compatible API key pool. Takes precedence over OPENAI_API_KEY and rotates to the next key on auth, quota, or rate-limit failures (* not needed for local models like Ollama, LM Studio, Atomic Chat, or other local OpenAI-compatible proxies). |
OPENAI_API_KEY |
Required only when OPENAI_API_KEYS is unset or empty for non-local OpenAI-compatible cloud routes* |
Your API key (* not needed for local models like Ollama, LM Studio, Atomic Chat, or other local OpenAI-compatible proxies). A comma-separated list also enables key rotation. |
OPENAI_MODEL |
OpenAI-compatible only | Model name such as gpt-4o, deepseek-v4-flash, or llama3.3:70b |
OPENAI_BASE_URL |
No | API endpoint, defaulting to https://api.openai.com/v1 |
OPENAI_API_BASE |
No | Compatibility alias for OPENAI_BASE_URL |
OPENCLAUDE_OLLAMA_NUM_CTX |
Ollama only | Request-level Ollama context window. Defaults to 32768; set a larger value for longer same-session history if your model and hardware can handle it. |
CLAUDE_CODE_OPENAI_CONTEXT_WINDOWS |
No | JSON map of OpenAI-compatible model names to context windows, such as {"custom-model":1000000}. Use this when a custom provider does not expose context metadata from /v1/models. |
CLAUDE_CODE_OPENAI_MAX_OUTPUT_TOKENS |
No | JSON map of OpenAI-compatible model names to max output tokens, such as {"custom-model":32768}. Use this when a custom provider does not expose output-limit metadata from /v1/models. |
OPENCODE_API_KEY |
OpenCode Zen / Go | Shared API key for OpenCode Zen (pay-as-you-go) and OpenCode Go (subscription); get yours from https://opencode.ai |
MIMO_API_KEY |
Xiaomi MiMo route | Xiaomi MiMo API key for https://api.xiaomimimo.com/v1; mirrored into the OpenAI-compatible auth env when the MiMo route is active |
CLAUDE_CODE_USE_GEMINI |
Gemini only | Set to 1 to enable the direct Gemini provider path |
GEMINI_API_KEY / GOOGLE_API_KEY |
Gemini API-key auth | Gemini API key for direct Gemini setup |
GEMINI_MODEL |
Gemini only | Model name such as gemini-3-flash-preview or gemini-2.5-pro |
GEMINI_BASE_URL |
No | Override the Gemini base URL |
CLAUDE_CODE_USE_MISTRAL |
Mistral only | Set to 1 to enable the dedicated Mistral provider path |
MISTRAL_API_KEY |
Mistral only | Mistral API key |
MISTRAL_MODEL |
Mistral only | Model name such as devstral-latest |
MISTRAL_BASE_URL |
No | Override the Mistral base URL |
CODEX_API_KEY |
Codex only | Codex or ChatGPT access token override |
CHATGPT_ACCOUNT_ID / CODEX_ACCOUNT_ID |
Codex only | Required for manual Codex env setup when the account id is not coming from auth.json or stored OAuth credentials |
CODEX_AUTH_JSON_PATH |
Codex only | Path to a Codex CLI auth.json file |
CODEX_HOME |
Codex only | Alternative Codex home directory |
OPENCLAUDE_MAX_RETRIES |
No | Maximum retry attempts for retryable API failures, capped at 100 (default: 10). Set to 0 to disable retries after the initial request. If unset, deprecated CLAUDE_CODE_MAX_RETRIES is still honored for compatibility. |
OPENCLAUDE_RETRY_DELAY_MS |
No | Base retry delay in milliseconds for APIs that do not send Retry-After; exponential backoff starts from this value, capped at 60000 (default: 500) |
OPENCLAUDE_QUERY_HARD_MAX_MS |
No | Foreground query hard maximum in milliseconds. Defaults to 1800000 (30 minutes). Use a larger positive integer for long autonomous sessions; invalid, zero, negative, fractional, or timer-overflow values are ignored with a warning. |
OPENCLAUDE_DISABLE_CO_AUTHORED_BY |
No | Suppress the default Co-Authored-By trailer in generated git commits |
OPENCLAUDE_LOG_TOKEN_USAGE |
No | When truthy (e.g. verbose), emits one JSON line on stderr per API request with input/output/cache tokens and the resolved provider. User-facing debug output — complements the REPL display controlled by /config showCacheStats. Distinct from CLAUDE_CODE_ENABLE_TOKEN_USAGE_ATTACHMENT, which is model-facing (injects context usage info into the prompt itself). Both can run together. |
Model env vars are provider-scoped: first-party Anthropic sessions read
ANTHROPIC_MODEL, OpenAI-compatible sessions read OPENAI_MODEL, Gemini reads
GEMINI_MODEL, and Mistral reads MISTRAL_MODEL. For manual Bedrock, Vertex,
or Foundry launches, select the model with --model.
Per-model limit overrides (settings.json)
When a custom OpenAI-compatible provider does not expose context metadata from
/v1/models, you can pin a model's context window and max output tokens. In
addition to the CLAUDE_CODE_OPENAI_CONTEXT_WINDOWS /
CLAUDE_CODE_OPENAI_MAX_OUTPUT_TOKENS env vars above, you can set a
modelLimits map in your settings.json (the same file /config writes, e.g.
~/.openclaude/settings.json):
{
"modelLimits": {
"my-custom-deployment": { "contextWindow": 262144, "maxOutputTokens": 32768 },
"api.private-llm.test:my-custom-deployment": { "contextWindow": 1000000 }
}
}
- Key matching — keys match the model api-name exactly, or by prefix (e.g.
my-custommatchesmy-custom-deployment-v2). An exact key always wins over a prefix key. A host-qualified key (<host>:<model>) only wins over a bare key within the same match kind — a host-qualified exact key beats a bare exact key, and a host-qualified prefix beats a bare prefix, but a bare exact key still beats a host-qualified prefix. So to give the same model different limits per endpoint, use host-qualified exact keys for each endpoint.<host>is theOPENAI_BASE_URLhost including the port when the URL has one (new URL(baseUrl).host): forhttp://localhost:4000/v1the key islocalhost:4000:my-model, notlocalhost:my-model. Either field may be omitted to override only one limit. - Precedence — from highest to lowest: an exact env-var override → the
built-in catalog value → the discovery-cache value → a prefix env-var
override →
modelLimits→ the descriptor default. (The built-in catalog is checked before the discovery cache.) So env-var overrides always win overmodelLimits, andmodelLimitsmainly fills in models that have no built-in metadata (a known catalog model keeps its catalog limit unless you set an exact env override for it).
Safety strictness
OpenClaude runs several "safety" checks: a model-level refusal directive, bash
command-injection validation, and sensitive-file / auto-edit guards. These are
conservative by design, but a few of them can surface as refusals or approval
prompts for entirely benign, routine coding tasks (e.g. editing .gitmodules,
running a build script that contains $(date), or writing a CTF port scanner).
See issue #1616.
Set OPENCLAUDE_SAFETY_LEVEL to dial strictness without changing behavior for
everyone:
| Value | Behavior |
|---|---|
strict |
Current/default-equivalent non-permissive behavior. |
balanced |
Default. Same behavior as strict. |
permissive |
Opt-in mode for users who prefer fewer false-positive stops. It bypasses the legacy bash command-injection validation path entirely, keeps ordinary interpreter allow-rules (Bash(python:*), Bash(npm run:*), …) when entering auto mode, and skips prompts for routine edits to filenames on the broad sensitive-file list. Dangerous directory, Windows-path, symlink-resolved path, and UNC guards remain active. The model-level prompt is not weakened by this flag. |
export OPENCLAUDE_SAFETY_LEVEL=permissive # relax benign-task false positives
Runtime Hardening
Use these commands to validate your setup and catch mistakes early:
# quick startup sanity check
bun run smoke
# validate provider env + reachability
bun run doctor:runtime
# print machine-readable runtime diagnostics
bun run doctor:runtime:json
# persist a diagnostics report to reports/doctor-runtime.json
bun run doctor:report
# print a redacted public issue report
openclaude doctor report --markdown
# write a redacted JSON issue report for attachment
openclaude doctor report --json --out openclaude-report.json
# write a deterministic task report from a session transcript
openclaude report --json --transcript ~/.openclaude/projects/-path-to-project/session-id.jsonl --out task-report.json
# print a human-readable task report from the latest session in the current project
openclaude report --markdown
# full local hardening check (smoke + runtime doctor)
bun run hardening:check
# strict hardening (includes project-wide typecheck)
bun run hardening:strict
Notes:
doctor:runtimefails fast ifCLAUDE_CODE_USE_OPENAI=1with a placeholder key or a missing key for non-local providers.doctor:runtimealso validates the dedicated Gemini and Mistral env paths whenCLAUDE_CODE_USE_GEMINI=1orCLAUDE_CODE_USE_MISTRAL=1.- Local providers such as
http://localhost:11434/v1,http://10.0.0.1:11434/v1, andhttp://127.0.0.1:1337/v1can run withoutOPENAI_API_KEY. - Codex profiles validate
CODEX_API_KEYor the Codex CLI auth file and probePOST /responsesinstead ofGET /models. openclaude doctor reportis redacted by default and is intended for GitHub issues. It summarizes provider/runtime/build/settings state without prompts, transcripts, raw settings files, API keys, MCP command details, or full home-directory paths.openclaude report --jsonandopenclaude report --markdownsummarize observed session facts such as tool uses, Bash commands, validation commands, changed files, branch metadata, warnings, and linked issue/PR references. Use--transcript <file>for an explicit transcript,--session <id>for a stored session, or omit both to report the latest session for the current project. Large previews are truncated and credential-shaped strings are redacted. When no validation command is observed, the report keepsvalidationsempty and includes a warning instead of claiming checks passed.
Provider Launch Profiles
Use profile launchers to avoid repeated environment setup:
# one-time profile bootstrap (prefer viable local Ollama, otherwise OpenAI)
bun run profile:init
# preview the best provider/model for your goal
bun run profile:recommend -- --goal coding --benchmark
# auto-apply the best available local/openai provider/model for your goal
bun run profile:auto -- --goal latency
# codex bootstrap (defaults to codexplan and ~/.codex/auth.json)
bun run profile:codex
# openai bootstrap with explicit key
bun run profile:init -- --provider openai --api-key sk-...
# gemini bootstrap with explicit key
bun run profile:init -- --provider gemini --api-key ...
# ollama bootstrap with custom model
bun run profile:init -- --provider ollama --model llama3.1:8b
# ollama bootstrap with intelligent model auto-selection
bun run profile:init -- --provider ollama --goal coding
# atomic-chat bootstrap (auto-detects running model)
bun run profile:init -- --provider atomic-chat
# codex bootstrap with a fast model alias
bun run profile:init -- --provider codex --model codexspark
# launch using persisted user-level provider profile
bun run dev:profile
# codex profile (uses CODEX_API_KEY or ~/.codex/auth.json)
bun run dev:codex
# OpenAI profile (uses the saved OpenAI profile, or OPENAI_API_KEYS / OPENAI_API_KEY from your shell)
bun run dev:openai
# Gemini profile (uses the saved Gemini profile, or GEMINI_API_KEY / GOOGLE_API_KEY from your shell)
bun run dev:gemini
# Ollama profile (defaults: localhost:11434, llama3.1:8b)
bun run dev:ollama
# Atomic Chat profile (Apple Silicon local LLMs at 127.0.0.1:1337)
bun run dev:atomic-chat
profile:recommend ranks installed Ollama models for latency, balanced, or coding, and profile:auto can persist the recommendation directly.
If no profile exists yet, dev:profile uses the same goal-aware defaults when picking the initial model.
Provider Profile Model Picker Mode
When a saved provider profile is active, /model can either show the provider's
catalog/discovered models or only the models explicitly listed in the profile.
Configure this in ~/.openclaude.json:
{
"providerProfileModelPickerMode": "auto"
}
Supported values:
auto(default): single-model profiles show the provider catalog; multi-model profiles show the explicit profile list; native vendor routes keep their full provider catalog.provider: show the provider catalog/discovery list first and append profile-only custom model IDs.profile: show only explicitly configured profile models.
When the provider-profile env workflow is active (i.e. a profile has been
applied and CLAUDE_CODE_PROVIDER_PROFILE_ENV_APPLIED=1 is set — as it is after
launching with a saved profile) and you have more than one saved provider
profile, /model also lists models from your inactive profiles, grouped
under their profile name. Selecting one activates that provider profile and
switches to the chosen model in a single step, reconciling fast mode if the
target provider cannot run it. These cross-profile entries appear only in the
interactive /model picker — they are never returned to SDK/automation callers
and are hidden from inline pickers (such as the prompt hotkey or Settings),
which cannot switch the active profile. Simply having multiple profiles
configured without the env workflow active does not surface them.
Use --provider ollama when you want a local-only path. Auto mode falls back to OpenAI when no viable local chat model is installed.
Use --provider atomic-chat when you want Atomic Chat as the local Apple Silicon provider.
Use profile:codex or --provider codex when you want the ChatGPT Codex backend.
dev:openai, dev:gemini, dev:ollama, dev:atomic-chat, and dev:codex
run doctor:runtime first and only launch the app if checks pass.
For dev:ollama, make sure Ollama is running locally before launch.
For dev:atomic-chat, make sure Atomic Chat is running with a model loaded before launch.
Message-Count Compaction Threshold
By default, OpenClaude compacts conversations based on token usage and also applies a safety hard cap of 1000 active messages. The hard cap catches long sessions that accumulate many small messages with negligible token cost.
This hard cap is a safety net: it can still trigger compaction even when
DISABLE_COMPACT, DISABLE_AUTO_COMPACT, or a disabled auto-compact setting
would otherwise prevent it. Set OPENCLAUDE_MAX_ACTIVE_MESSAGES_HARD_CAP=0
only when you need to suppress that safety cap for diagnostics.
If you frequently resume long sessions that accumulate hundreds of small
tool-result messages with negligible token cost, you can opt in to message-count
compaction via the in-app /config command:
/config
Select Message-count compaction and choose a threshold (100, 200, 500,
or 1000). Setting it to off (default) leaves only the built-in hard cap.
This setting is intended for power users debugging specific edge cases. Most
users should leave it at off.
The legacy OPENCLAUDE_MAX_ACTIVE_MESSAGES environment variable is still
honored when the setting is off. OPENCLAUDE_MAX_ACTIVE_MESSAGES_HARD_CAP
can override the safety cap; set it to 0 only for diagnostics.
Long-session memory guard validation
For changes that touch auto-compact, provider request conversion, transcript retention, or in-process teammates, run the focused long-session guard checks:
bun test --feature=UNATTENDED_RETRY src/query/autoCompactCooldown.test.ts src/utils/maxActiveMessages.test.ts src/services/api/openaiShim.test.ts
These tests cover repeated over-cap turns, auto-compact cooldown blocking, teammate active-message compaction, malformed hard-cap overrides, and pruned-history tool-call/tool-result pairing. They are not a substitute for a multi-hour manual soak, but they pin the bounded-history and conversion invariants that previously let long sessions grow until Node/V8 OOM.