Commit Graph
1170 Commits
Author SHA1 Message Date
BogdanandGitHub 09eba26d30 feat(cost): support exact custom model pricing (#2131)
* feat(cost): support exact custom model pricing

* fix(cost): address custom pricing review feedback
2026-08-16 15:55:02 +08:00
BogdanandGitHub c30578819e diagnostics(query): trace interruption causality (#2111)
* diagnostics(issue-1830): trace interruption causality

* test(issue-1830): lock interruption ownership matrix

* fix(codex): preserve stream deadline contract

* fix(diagnostics): harden interruption trace lifecycle

Refs #1830

* fix(diagnostics): harden interruption trace settlement

Refs #1830

* fix(diagnostics): preserve interruption causality

* fix(diagnostics): address interruption trace review

* fix(diagnostics): preserve tracing observer contracts

* fix(diagnostics): preserve interruption trace contracts

* test(permissions): cover interactive hook interrupts
2026-08-16 15:54:29 +08:00
BogdanandGitHub ea655163d3 feat(zai): expand Coding Plan catalog support (#2127)
* feat(zai): expand Coding Plan catalog support

Signed-off-by: chioarub <chioarub@gmail.com>

* fix(zai): use supported low reasoning mode

Signed-off-by: chioarub <chioarub@gmail.com>

---------

Signed-off-by: chioarub <chioarub@gmail.com>
2026-08-15 17:05:49 +08:00
0xfandomandGitHub 6c7a12b2a2 fix(websearch): reject non-positive WEB_CUSTOM env overrides (#2124)
The custom web-search provider read WEB_CUSTOM_TIMEOUT_SEC and
WEB_CUSTOM_MAX_BODY_KB with Number(x) || DEFAULT. That idiom only
rescues 0 and NaN — a negative or Infinity value is truthy and passes
straight through. A negative WEB_CUSTOM_MAX_BODY_KB makes the
"body exceeds N bytes" check reject every POST search, and a negative
WEB_CUSTOM_TIMEOUT_SEC drives an immediate abort on every request.

Route both through readPositiveEnvNumber, which falls back to the
default for missing, empty, non-finite, or non-positive input.
2026-08-14 10:13:27 +08:00
0xfandomandGitHub f553d0896d fix(api): resolve swarm-field tool names by own-property (#2123)
filterSwarmFieldsFromSchema() looked the tool name up with a bare
SWARM_FIELDS_BY_TOOL[toolName] on a plain-object map. A tool whose name
collides with an Object.prototype member (constructor, hasOwnProperty,
isPrototypeOf, propertyIsEnumerable) resolves the inherited function,
which is truthy with .length === 1 — so the empty/undefined guard is
bypassed and the subsequent for...of throws "is not iterable", failing
schema construction for the entire request.

Guard the lookup with Object.hasOwn so a prototype-named tool is treated
as unmapped, matching the own-key guard already used for provider-
supplied tool names in services/api/toolArgumentNormalization.ts.
2026-08-14 10:12:53 +08:00
fb9102c422 feat(aimlapi): passwordless onboarding and resumable card top-up (3/3) (#2032)
* feat(aimlapi): add checkout state persistence and sign-in key cache

* fix(aimlapi): cover lock recovery and complete the reset receipt

* test(aimlapi): cover CAS result, reset receipt and sign-in key permissions

* fix(aimlapi): make stale-lock recovery ownership-safe across processes

* test(aimlapi): pass a file URL specifier to the lock workers

* fix(aimlapi): use proper-lockfile and harden checkout-state persistence

* fix(aimlapi): preserve issued keys and survive stale-lock steal races

* fix(aimlapi): surface swallowed lock-retry conditions

* feat(aimlapi): resume interrupted checkouts in the top-up entry points

* fix(aimlapi): keep resumable checkouts through transient and paid states

* fix(aimlapi): surface a lost settled-receipt write to the caller

* fix(aimlapi): harden checkout resume across CLI resume, idempotency and transient errors

* fix(aimlapi): resume settling checkouts and preserve records on ambiguous reads

- Treat 'exchanging' as a paid/resumable status so a run interrupted between
  payment and receipt resumes the exchange instead of opening a second,
  chargeable checkout (matches pollUntilPaid).
- Preserve the recorded checkout on a malformed-but-successful status read
  (AimlapiApiError status 200), alongside the existing transient-error path.
- Print the full recovery key in the receipt-write-failed warning; a masked
  key is useless as a last copy.
- Re-register the real client/providerProfile modules in afterAll so the
  test stubs cannot leak into later files (mock.restore does not undo
  mock.module).
- Bound each lock worker's exit before draining its pipes, and loosen the
  held-lock timeout assertion to any errno (Windows is not always ELOCKED).

* fix(aimlapi): stop mock leak, fail closed on spent sessions, clear GUI receipt

- topup.test.ts: capture the real client/providerProfile modules through a
  cache-busting query so afterAll restores the genuine module instead of the
  corrupted (stub-mutated) reference. Without this the './client.js' stub bled
  past afterAll and failed 25 client.test.ts cases whenever it ran first.
- resolveCheckoutSession: fail closed when a resumed session is already
  'exchanged' but no settled receipt survived locally, matching pollUntilPaid,
  instead of opening a second chargeable checkout for a one-shot key that is
  already gone.
- provisionAimlapiKey: return a clearReceipt closure; ProviderManager now calls
  it only after persistDraft actually saves, so a second GUI top-up opens a
  fresh checkout instead of short-circuiting to a stale key or throwing.
- claimAimlapiTopupState: refuse to replace a stored record that still has an
  open resume token for a different intent, so a changed amount cannot strand a
  still-payable checkout and open a second one.
- Mask the issued key in the CLI receipt-write-failed warning; a paid-for
  credential should not land in scrollback.
- index.ts: note the live CLI/GUI callers and the clear-receipt obligation.
- Regression coverage for exchanged fail-closed, the claim guard, and the GUI
  receipt clear.

* fix(aimlapi): inject the topup transport instead of mocking client.js globally

The prior mock.module('./client.js') stub in topup.test.ts leaked past its
afterAll into client.test.ts (25 failures when this file ran first in CI,
bun 1.3.13). Replace it with a local injection seam:

- topup.ts exposes setAimlapiTopupTestDoubles, and both entry points create
  their client / write their profile through it (defaults unchanged, so
  production behaviour is identical).
- topup.test.ts injects a stub transport through that seam and no longer calls
  mock.module at all, so nothing can bleed OUT to client.test.ts.
- It still loads topup.js through a cache-busting ?ts= query so it stays immune
  to ProviderManager.test.tsx's mock.module('../integrations/aimlapi/index.js'),
  which mock.restore() does not undo and which would otherwise replace the
  shared provisionAimlapiKey binding (verified: without the query the barrel
  stub reaches this file and every topup test fails).

* fix(aimlapi): converge racing checkouts and preserve state on ambiguous reads

- Close a post-claim race: two runs of the same intent converge on one payment
  id, then each can open a session before either records one, and the second
  save overwrote the first's resume token (two payable checkouts). Add
  recordAimlapiCheckoutSession, a compare-and-swap that only records while the
  token is empty; resolveCheckoutSession adopts the winner's session and
  abandons the one it just opened, so pay() converges idempotently on a single
  charge.
- Resume path now preserves the record on any ambiguous getSession error
  (transient, malformed-200, auth/4xx) and retires it only on a definitive
  404/410 gone-session, instead of clearing on every non-transient error.
- pollUntilPaid retries a malformed-but-successful (status 200) body instead of
  aborting, matching the resume path.
- Regression coverage for the peer-records-first race, ambiguous-error
  preserve, 404 replace, and poll-retry-on-200.

* fix(aimlapi): re-validate an adopted peer session before paying it

When two racing runs converge and this one adopts the peer's recorded session,
route that session through the same status classification as the initial
resume: return it only while resumable, fail closed on 'exchanged', and surface
a re-run error on any other terminal status instead of calling pay() on a dead
session. Add regression coverage for an adopted session that is exchanged or
cancelled in the race window.

* test(aimlapi): make abandoned-lock recovery test deterministic

The stale-lock recovery test asserted that all racing claims return the same
payment id. Under a stale-lock steal two recoverers can briefly hold the lock
and mint distinct ids, so that assertion was flaky in CI. A diverged claim is
harmless — it is refused at its next compare-and-swap — so the invariant that
actually matters is that exactly ONE checkout gets established. Drive the
workers through the full claim->record flow and assert a single winner, which
the single-slot store plus the record CAS guarantee deterministically.

* fix(aimlapi): acquire checkout-state off the interactive thread and recover fresh orphans

The interactive top-up flow acquired the checkout-state lock synchronously,
parking the Ink event loop (UI, timers, SIGINT) on Atomics.wait for up to the
5s timeout, and that timeout was shorter than the 30s stale window so a lock
orphaned by an interrupted holder could not be recovered on an immediate resume.

- Add withStateLockAsync + async variants of the state mutators, sharing the
  same inner operations. It acquires with the timer-free lockSync (a sub-ms
  mkdir) but yields via await between retries, so the UI stays live, and its
  longer deadline (15s) covers the stale window so a fresh orphan is reclaimed
  once stale rather than timing out. Shrink the stale window to 8s (sub-ms
  sections never approach it) so recovery is quick.
- Route topup.ts (CLI + GUI) through the async variants; provisioned.clearReceipt
  is now async, and ProviderManager awaits it best-effort so a cleanup failure
  cannot surface an error that invites a retry (a duplicate provider profile).
- Regression coverage for async orphan recovery (mutation-checked: a deadline
  below the stale window fails to recover). Generous per-file test timeout since
  the now-async provisioning yields to a loaded runner's event loop.

* test(aimlapi): cover the async state mutators and speed up the orphan-recovery test

- Add thin contract tests for saveAimlapiTopupStateAsync (CAS accepted/rejected),
  recordAimlapiCheckoutSessionAsync (compare-and-swap on the empty resume token),
  and clearAimlapiTopupStateAsync (ownership-scoped clear); the CLI/GUI flow now
  routes through these and only claimAsync was exercised.
- Back-date the orphaned lock to just inside the stale window so the async
  recovery test reclaims it in ~2s instead of burning the full 8s window.

* fix(aimlapi): fail closed on corrupt state, deliver keys on receipt failure, add a reset action

- Fail closed when the checkout-state file is present but unreadable/schema-
  invalid instead of reading it as absent: a claim would otherwise overwrite an
  open/paid checkout or an exchanged key and open a second chargeable one.
- Never strand a one-shot exchanged key: a receipt-write that throws (lock
  timeout / fs / corrupt state), not just a lost CAS, is caught so the CLI still
  writes the profile and the GUI still returns the key; the post-delivery clear
  is best-effort too.
- Add an explicit discard/reset escape hatch — discardAimlapiCheckoutState, an
  "openclaude aimlapi reset" command, and a GUI "Start over" on the top-up error
  screen — so a terminal checkout (whose resume token blocks a different intent)
  or a corrupt state file can be cleared without editing internal files.
- Surface a failed GUI receipt retirement: retry a few times, then show a
  non-blocking warning instead of swallowing it silently.
- Only chmod a config directory this flow actually created (via mkdirSync's
  return); never tighten a pre-existing OPENCLAUDE_CONFIG_DIR root. The state
  file's own 0600 mode protects the credential.
- Tests for each, incl. corrupt/schema-invalid fail-closed, receipt-write-throw
  key delivery (CLI + GUI), discard semantics, retried/surfaced GUI cleanup, and
  a POSIX check that an existing config dir keeps its mode.

* fix(aimlapi): surface the CLI recovery-receipt clear failure and cover the reset handler

- The CLI finishCliTopup clear failure only went to logForDebugging (debug-only),
  invisible to the user, unlike the loud receipt-write warning and the GUI
  warning. Print a [warn] line pointing at "openclaude aimlapi reset" so a
  stranded receipt that blocks a different top-up is not hidden.
- Add tests for the aimlapiReset CLI handler (discards a stored checkout /
  reports when there is nothing to discard).
- Assert the CLI clear-failure warning is surfaced in the receipt-write-failure
  test (mutation-checked).
- Document "openclaude aimlapi reset" and the GUI Start over recovery in
  docs/aimlapi-setup.md.

* fix(aimlapi): protect a settled receipt from reset and bind method into the checkout identity

- Discard (CLI reset / GUI Start over) no longer deletes a settled receipt — the
  only copy of a paid-for, one-shot key — unless forced. discard now returns
  'discarded' | 'kept-settled' | 'none'; the CLI adds --force and the GUI refuses
  and points back at the recovering retry.
- Add the payment method to AimlapiTopupIntent so a card->crypto (or reverse)
  restart is a different intent and cannot adopt the prior checkout's reused
  idempotency id on the wrong rail; covered by a changed-method resume test.
- Narrow the sign-in-key cache: it is a persistence primitive for the follow-up
  guided passwordless flow with no in-tree consumer, so drop it from the public
  barrel and document the scope (kept in topupState.ts for that follow-up).
- Regression + mutation coverage for the settled-receipt protection and the
  method-scoped intent.

* fix(aimlapi): back off receipt-clear retries, expand the discard API, and test Start over

- Add a 150ms backoff between the GUI clearReceipt() retries so the loop can
  actually ride out a lock-contention window instead of exhausting three
  back-to-back attempts.
- Re-export AimlapiDiscardResult and add resetAimlapiCheckoutSessionAsync so
  barrel consumers can name the discard outcome and use a non-blocking reset,
  matching the other mutators.
- Cover the GUI Start over recovery: it is 'r' (settings:retry) on the top-up
  error screen — the Settings context has no confirm:yes and Enter closes the
  panel; tests drive both the discard and the kept-settled refusal path.
- Rename the test's stale-age constant (LOCK_AGE_WELL_PAST_STALE_MS) so it no
  longer reads as the source's 8s window.

* fix(aimlapi): serialize the one-shot key exchange and recover receipts before login

Address three checkout-state review findings:

- [P1] Serialize the non-idempotent key exchange behind an exchange lease so
  racing same-intent processes mint and record the credential exactly once. The
  elected lease holder exchanges and records the settled receipt; a peer that
  loses the election waits for that receipt and resumes from it instead of
  exchanging in parallel; a lease abandoned by a crashed holder goes stale (past
  the client request timeout) and is reclaimed on a later attempt. The lease is
  released on a failed exchange so a retry is not blocked for the stale window.

- [P1] Recover a settled local receipt BEFORE authenticating in both the CLI and
  guided entry points. A run interrupted after the one-shot exchange but before
  the profile write leaves the paid-for key only in that receipt; requiring a
  fresh login to reach it stranded the key whenever the password had changed or
  the auth service was down. The receipt read is side-effect free and needs no
  token, so it now runs first and only authenticates when a checkout must be
  created/resumed/exchanged.

- [P2] Fix the guided-recovery key in the setup guide: Start over is bound to r,
  not Enter (Enter closes the settings panel).

Covered by state-layer lease tests (acquire/held/stale-steal/settled/gone/
release), a two-process race asserting the exchange runs exactly once with the
loser resuming the receipt, and no-auth-on-settled tests for both entry points.

* test(aimlapi): harden the exchange-lease coverage

Address review follow-ups on the exchange-lease tests (all test-only):

- Gate the two-process race on the loser's own "waiting" status signal instead
  of a fixed sleep, and assert it fired, so the test deterministically exercises
  the held -> wait -> resume path rather than possibly reading settled directly.
- Type seedPersistedState's overrides as Partial<AimlapiPersistedTopup> so a
  misspelled/wrong-typed seed key is a typecheck failure instead of a record
  that silently reroutes the test down another branch.
- Cover re-acquiring your own lease (leaseOwner === owner) so the guard that
  keeps a caller from mistaking its own fresh lease for a live peer's — and
  self-blocking until the stale window — cannot regress unnoticed.

* fix(aimlapi): fence a superseded profile write, protect unreadable state, resume the receipt model

Address three checkout-state review findings:

- [P1] Fence an in-flight checkout before it writes its key to the provider
  profile. If a reset (or a fresh top-up) replaces the stored slot while an
  abandoned flow is awaiting client.exchange(), that flow's settled-receipt CAS
  now misses; previously it still went on to write its stale key and could
  clobber the profile the new top-up created. recordSettledReceipt now reports
  recorded | superseded | errored, and the exchange path aborts (rejecting, and
  pointing the user at rotating the orphaned key) on `superseded` while still
  delivering on a transient `errored` so a paid-for key is never stranded.

- [P1] Do not discard an unreadable/corrupt state file without --force. Such a
  file could be a damaged receipt holding the sole copy of a one-shot key, so
  discardAimlapiCheckoutState now returns `kept-unreadable` and keeps it unless
  forced — matching the safety promise (and existing settled-receipt protection)
  that reset never loses an issued key. The CLI and guided "Start over" surface
  the new outcome and point at `reset --force`; docs updated.

- [P2] Use the settled receipt's model when a peer completed the exchange. The
  settled lease branch dropped lease.state.model, so a loser resuming another
  run's receipt configured its own --model instead of the one actually
  provisioned; it now propagates the receipt's model through both callers.

Covered by: a superseded-mid-exchange fence test, corrupt-discard-needs-force
tests (state layer + CLI + guided GUI), and a two-process race asserting the
loser adopts the winner's provisioned model. All three are mutation-proven.

* test(aimlapi): drop the flaky third guided Start-over drive test

The guided "Start over on a kept-unreadable discard" test added a third
consecutive full Ink mount+drive to ProviderManager.test.tsx. On the CI-pinned
bun (1.3.13) that destabilises the Ink stdin harness (`stdin.ref is not a
function`), so the 'r' keypress never reaches the handler and the awaited
discard call never fires — a timeout unrelated to the code under test (it passes
on local bun 1.3.14). The two-drive configuration is green on CI.

The kept-unreadable behaviour stays covered where it is deterministic: the state
layer (kept-unreadable unforced, discarded on --force) and the CLI handler (the
`reset --force` guidance). The guided keybinding→discard→refusal plumbing is
covered by the identical-structure kept-settled drive test; the kept-unreadable
GUI branch is a trivial mirror of it. A comment records why the third drive is
intentionally omitted.

* wip(aimlapi): converge integration + CLI to the passwordless card-only flow (#1988)

Mid-port checkpoint. Rewrites the AI/ML API integration backend and CLI to the
canonical passwordless, card-only design (target PR #1988), in the current
code style. NOTE: the branch does not yet compile — the GUI (ProviderManager
passwordless rewire) and the aimlapi test suite still reference the removed
API and are the remaining work.

Done (typecheck-clean in these files):
- config.ts: add payBaseUrl + verificationBaseUrl endpoints, buildPartnerReturnUrl.
- topupState.ts: reduce to the #1988 shape — intent keyed on payBaseUrl/
  verificationBaseUrl (no `method`); sync lockfile-based state; drop the
  exchange-lease, discard/reset-command surface, async variants and
  fail-closed-on-corrupt (corrupt reads as absent).
- client.ts: drop the password path (signup/login) and PaymentMethod/crypto;
  pay() is card-only.
- topup.ts: rewrite to the passwordless phase-machine flow (checkAccount ->
  code sign-in / new-account -> provision or by-key top-up; resolveTopupSession;
  pollUntilPaid / pollUntilExchangeSettled / pollUntilByKeyToppedUp). Keeps the
  DI test seam (no cross-file mock.module). Profile env writes AIMLAPI_API_KEY
  mirror + CLAUDE_CODE_PROVIDER_ROUTE_ID.
- onboarding.ts, messages.ts (canonical copy), validation.ts, index.ts (lean
  barrel), providerManagerAimlapi.ts (GUI indirection layer): new.
- CLI: aimlapiCommand.ts (registerAimlapiCommand, --email/--code/--amount, no
  --method/reset), main.tsx wiring, handlers/aimlapi.ts (redacted errors).

Mandatory attribution headers wired on EVERY aimlapi request: X-AIMLAPI-Source
(agent/openclaude) + X-AIMLAPI-Partner-ID — in client.ts request() (auth/
checkout) and the config attribution path (inference/catalog).

* wip(aimlapi): port the passwordless provider-manager GUI (#1988)

Rebase ProviderManager.tsx on the #1988 passwordless aimlapi flow (email ->
6-digit code -> low-balance -> top-up / paste-existing-key -> done), replacing
the old password/method/Start-over flow and importing the aimlapi surface from
providerManagerAimlapi.js. Re-apply the current newer-main, non-aimlapi
`apiFormat: 'auto'` feature that the rebase would otherwise revert (form
metadata, toDraft default, display label, startCreateFromPreset default, the
API-format picker's Automatic option, and persistDraft's selectedApiFormat
'auto' -> undefined branch, kept alongside #1988's deferNavigation/onSaved).
Re-export resolveRouteCredentialValue from integrations/index for the GUI.

All source now typechecks; the aimlapi test suite is the remaining work.

* test(aimlapi): port integration + CLI tests to the passwordless flow

- topupState/topup/onboarding/aimlapiCommand tests: port #1988's coverage,
  adapting the transport to `globalThis.fetch` stubbing and the profile/prompt
  doubles to the module's `setAimlapiTopupTestDoubles` DI seam (no process-global
  mock.module, which leaks across files in this repo).
- client/config tests: keep the current repo's stricter versions (complete
  session contracts), pruning the removed password/crypto cases.
- Add mandatory-attribution-header coverage on all four request classes: the
  client sends X-AIMLAPI-Source + X-AIMLAPI-Partner-ID on auth/checkout, and the
  config attribution path sends both on inference/catalog and strips them for a
  non-canonical proxy endpoint.
- Remove the reset-based CLI handler test (reset no longer exists).

Full aimlapi integration + CLI suite green (67 tests). ProviderManager GUI test
is the remaining piece.

* test(aimlapi): port the passwordless provider-manager GUI tests (#1988)

Rebase ProviderManager.test.tsx on #1988's version for the passwordless aimlapi
GUI tests (email -> code -> low-balance -> top-up / paste-existing-key), which
mock ./providerManagerAimlapi.js. Re-apply HEAD's newer-main apiFormat 'auto'
test cases (API-mode picker, token field, OpenAI/GPT-5/MiniMax presets) since
the source keeps that feature, and restore the current preset list ('LongCat')
in the test's PRESET_ORDER so navigateToPreset indexes match the real presets.

ProviderManager suite green (42 tests).

* docs(aimlapi): rewrite the setup guide for the passwordless card-only flow (#1988)

Describe the passwordless /provider flow (saved-key continue, new-user email +
6-digit code, paste-existing-key, low-balance top-up) and the card-only CLI
`aimlapi topup --email/--code/--amount` (no --method, no reset). Document the
full endpoint override set and the two mandatory attribution headers
(X-AIMLAPI-Source + X-AIMLAPI-Partner-ID) sent on every request, stripped for a
non-canonical proxy endpoint.

* refactor(aimlapi): let tests inject prompt doubles into the top-up flow

* fix(aimlapi): lock the partner id and complete mandatory-header coverage

Lock the partner id to OpenClaude's own attribution id: drop the --partner-id
CLI flag and the AIMLAPI_PARTNER_ID env override so rebate/revenue-share
attribution can never be redirected. resolvePartnerId() now always returns the
built-in id; the mandatory X-AIMLAPI-Partner-ID header itself is unchanged.

Assert the mandatory X-AIMLAPI-Source header on the catalog/discovery and
inference paths (discoveryService, bootstrap, runtimeMetadata) — the header was
already sent, only the test expectations lagged. Refresh stale password-era copy
in the interactive prompt and a top-up comment left over from the removed flow.

* fix(aimlapi): address CodeRabbit review — settlement, key-safety, redaction

- Wait for a resumed sign-in top-up to settle before returning: the account
  non-exchange path now mirrors the by-key flow, so a credited balance is never
  reported while the billing operation is still in flight.
- Preserve a freshly minted sign-in key when the balance read is aborted, so an
  abort cannot orphan a paid credential and mint a second key on the next run.
- Clear the first-run env-key adoption markers when validation fails, so a retry
  re-validates instead of short-circuiting into persisting the unvalidated key.
- Never crash the top-up success screen on an amount parse edge — fall back to
  the raw entered amount after the payment has already cleared.
- Show all four API-format options (visibleOptionCount 3 -> 4).
- Document that guided provisioning requires the canonical inference endpoint.
- Add regression tests: resumed-sign-in settlement, aborted-balance key
  retention, failed-env-key re-validation, and CLI error redaction.

* test(aimlapi): wait for the masked code frame instead of a fixed delay

The AIMLAPI code-screen assertion captured output after a fixed 25ms sleep,
which is too short on a slower CI runner (Node 24) and intermittently missed the
freshly rendered mask characters. Wait for the masked frame instead so the
assertion is deterministic.

* test(aimlapi): harden the provider-manager GUI flows against CI timing

The GUI top-up flow intermittently failed on a loaded CI runner: the final
keystroke on the success screen was sent in the same tick as the render, before
Ink attached the input handler, so it was dropped and the flow stranded on the
done screen. Add the same input settle the other steps already use.

Also raise the shared waitForCondition default timeout (2s -> 5s). The predicate
is polled every 10ms and returns as soon as it is met, so this only adds patience
for a slow runner and never slows a passing wait — keeping the Ink-driven flows
deterministic under CI load.

* fix(aimlapi): gate attribution headers by trusted AI/ML API host

The client sent the mandatory source/partner headers on every request, but the
auth/app/pay/inference base URLs are all env-overridable — so a request pointed
at a user proxy (notably the balance probe against an overridden inference URL)
leaked OpenClaude's partner/source identity. Send them only when the resolved
request host is aimlapi.com (production or staging, over HTTPS), mirroring the
inference/catalog stripping contract in resolveAimlapiAttributionHeaders. Adds
an isTrustedAimlapiRequestUrl predicate plus canonical-sends / proxy-withholds
regression tests.

* test(aimlapi): wait for the done screen to settle before the final keystroke

Replace the fixed 25ms delay before the success-screen keystroke with an
observable frame-stability wait, so a slow CI runner cannot drop the keystroke
before Ink has committed the render and attached its input handler.

* fix(aimlapi): durable receipts and atomic election for concurrent top-ups

- Restore the atomic checkout-session election dropped during the passwordless
  convergence: recordAimlapiCheckoutSession is a first-writer-wins CAS, so two
  concurrent runs of the same intent settle on ONE payable checkout — a loser
  adopts the winner's token and abandons the session it just opened instead of
  leaving two chargeable checkouts. Wired through resolveTopupSession (the create
  branch elects then adopts; the resume branch notifies once) and both the CLI
  and GUI onSession callbacks.
- Persist the settled receipt (apiKey / apiKeyId / model / settled) in the GUI
  BEFORE the profile write, so an interrupted or failed write resumes with the
  paid, one-shot exchanged key instead of stranding it (mirrors the CLI).
- Clear the sign-in key cache with the just-minted key id on a sufficient-balance
  sign-in: persistDraft runs onSaved synchronously, so the aimlapiIssuedKeyId
  state setter has not applied yet — pass the id explicitly.

Adds regression tests for the election (first-writer-wins + loser adoption), the
settled-receipt ordering, and the sufficient-balance cache clear.

* fix(aimlapi): abort on a lost election and keep the receipt write best-effort

- Treat a null recordAimlapiCheckoutSession result as "the slot was cleared by a
  sibling that already completed this top-up" and abort, instead of silently
  proceeding to pay a second, unrecorded checkout (both the CLI persistSession
  and the GUI reportSession). Closes the residual double-charge race.
- Make the GUI settled-receipt write best-effort: the payment already cleared, so
  a receipt-write failure (lock contention, full/read-only disk) must not divert
  the flow into the top-up error path — the profile write is what matters.
- Align the recordAimlapiCheckoutSession test double with the real semantics:
  match on intent + payment id only and return null on a non-matching slot.

Adds a regression test that a sibling clearing the checkout mid-flow aborts
before any /pay call. The sufficient-balance sign-in test now waits for the code
screen to settle before typing (the transition dropped the first keystroke).

* test(aimlapi): settle after each awaited frame so keystrokes aren't dropped

The provider-manager GUI tests type on the line after waitForFrameOutput matches
a new screen, but Ink registers input handlers in an effect that runs after the
render commits. On a loaded CI runner the first post-transition keystroke could
be dropped, stranding the flow and timing out (seen intermittently on Node 22).
Add a short settle after every frame match — returning the same matched frame,
so no assertion changes — which lets the input handler attach before the caller
types. Fixes the class instead of patching individual call sites.

* fix(aimlapi): recover a settled GUI receipt and harden checkout/key edge cases

- Recover a settled checkout receipt in the provider-manager GUI before
  provisioning: if a prior run paid + exchanged the key and saved the receipt but
  was interrupted before the profile write, finish that write with the retained
  key instead of re-entering provisioning against the now-exchanged session
  (which fails in resolveTopupSession and strands the paid credential). Mirrors
  the CLI.
- Reject a non-HTTPS checkout payUrl at the client response boundary (the
  validator required only "openable"), so a session is never retained with an
  address the flow refuses later and then polls with no usable link.
- Do not discard a freshly minted sign-in key when its cache write fails: copy
  the key into memory before persisting and make the sign-in-cache / top-up-state
  writes best-effort, in both the GUI and the CLI, so a lock/permission/disk
  failure can't force a second key on retry.
- Narrow the setup guide: the canonical-inference requirement applies to
  new-account onboarding + key provisioning; the existing-key top-up runs against
  the configured endpoint.

Adds regression tests for the settled-receipt recovery (no re-provision) and the
non-HTTPS payUrl rejection.

* fix(aimlapi): HTTPS checkout callbacks, per-email key cache, safer edges

- Require a credential-free HTTPS base for the checkout return URLs (they embed
  the resumable session token) and for the browser return/landing URL, so a
  cleartext AIMLAPI_PAY_URL/AIMLAPI_RETURN_URL override can't leak the token or
  break the documented HTTPS return-target contract.
- Store sign-in recovery keys as an email-keyed collection instead of a single
  global record, so a concurrent/interrupted sign-in for one account no longer
  evicts another's key (which forced a duplicate mint). Old single-record files
  migrate on read; clear stays per-email ownership-aware.
- Treat post-success receipt cleanup as best-effort in both the CLI (finishProfile)
  and the GUI (resetAimlapiCheckoutIntent): the profile is already saved, so a
  lock/permission/IO failure clearing the receipt must not report failure.
- Reject scientific-notation amounts: parseAimlapiAmountUsd now requires a plain
  decimal with at most two fractional digits, closing the "20.001e0" sub-cent
  bypass that silently rounded to a wrong charge.
- Add the pay/verification/return env vars to the config-test snapshot so a set
  override can't pollute default-endpoint assertions.

Adds regression tests for each.

* fix(aimlapi): async checkout-state clear for the Ink flow + reject malformed bases

- Clear the checkout receipt through an async lock in the provider-manager GUI:
  restore withStateLockAsync + clearAimlapiTopupStateAsync and fire it
  best-effort (unawaited) from the save callback, so a contended lock no longer
  blocks Ink input/timers/SIGINT after the profile is already saved. The CLI keeps
  the sync clear (one-shot command).
- Reject a query string or fragment in the checkout/return base URLs: a base like
  https://pay.aimlapi.com/#x would swallow the appended /checkout?...sessionToken
  into the fragment, so the token never reaches the callback as a query param.
- Surface a non-fatal CLI note when receipt cleanup fails (the profile is already
  saved and the stale receipt reconciles on the next run).

Adds regression tests: async clear ownership, query/fragment rejection, and the
legacy single-record sign-in-cache migration.

* fix(aimlapi): reject bare ?/# delimiters in checkout and return base URLs

url.search / url.hash are empty for a bare delimiter (e.g. https://pay.aimlapi.com/?
or .../#), so those slipped past the query/fragment guard and still corrupted the
appended /checkout?...sessionToken=... . Reject any raw ?/# in the candidate in
both safeHttpsBaseUrl and requireHttpsBaseUrl, and cover the bare-delimiter cases.

* fix(aimlapi): harden checkout recovery — exchange lease, retry modes, payable guard

Addresses a fresh review round on the checkout state machine:

- Restore the cross-process exchange lease (dropped in the passwordless
  convergence): the one-shot key exchange is serialized so two processes resuming
  the same paid sign-up session cannot both exchange and strand the credential —
  the lease winner exchanges, peers wait for its settled receipt.
- Persist the exchange mode in the receipt so a retry that has since become
  sign-in still exchanges the paid session instead of minting an unrelated key
  and clearing the paid checkout (CLI + GUI).
- Recover the checkout URL on a pending_payment resume by re-issuing the
  idempotent pay/top-up (the stable paymentSessionId prevents a double charge)
  instead of polling a session the user can never open.
- Route settled-receipt recovery through persistExistingAimlapi for an existing
  saved profile / AIMLAPI_API_KEY top-up, so it updates the selected profile
  (preserveEnv) rather than minting a new one and copying the env key.
- Confirm before abandoning an already-open checkout: editing amount/auto-top-up
  after a checkout URL was opened now requires an explicit re-submit (the old
  browser tab stays chargeable and no endpoint can cancel it).
- Treat a credentials/query/fragment inference base as non-canonical so a
  `.../v1#x` override cannot be written as OPENAI_BASE_URL and break the shim.

Tests: exchange-lease election + failed-exchange release, retry-exchanges-the-
paid-session, idempotent URL recovery on resume, canonical-gate rejection, and
the re-edit confirmation.

* fix(aimlapi): repair exchange-lease liveness and the re-edit abandon guard

- exchange lease: a peer that finds a live foreign lease now re-attempts on each
  poll instead of only watching for a settled receipt, so it resumes the moment
  the holder settles OR frees the lease (failed/crashed) rather than hanging the
  full 20-minute poll window; folds the wait into the lease loop.
- exchange lease: treat a future-dated exchangeLeaseAt (backwards clock jump or
  an edited state file) as stale and reclaim it, instead of reading a negative
  age as perpetually fresh and deadlocking every peer.
- re-edit guard: reset the abandon acknowledgement when a new checkout opens so a
  further edit to a different amount/auto-top-up is confirmed again instead of
  silently abandoning the freshly-opened chargeable tab; clear the opened-checkout
  tracking once payment settles so a later re-edit never warns about a paid tab.
- tests: lease release is owner-scoped and preserves a settled receipt; a
  future-dated lease is reclaimed; the GUI re-edit warning re-arms after a second
  edit.

* test(aimlapi): sync re-edit test on rendered amount; guard vacuous lease seed

- re-edit GUI test: submit only once the edited amount is reflected in the
  rendered frame instead of after a fixed 25ms delay, so Enter is never processed
  against the stale amount on a slow runner.
- future-dated lease test: assert the seed compare-and-swap actually persisted the
  lease before acquiring, so the reclaim path can never pass vacuously.
- exchange lease: record a swallowed release failure via file-backed debug logging
  (safe on the Ink GUI path) so a lock/permission problem behind a slow takeover is
  diagnosable.

* test(aimlapi): match the complete edited amount in the re-edit frame wait

Prefix matching let "$250" match a stray "$2500" (and "$2500" match "$25000"),
so a wrong-amount input regression could pass unnoticed. Pin the complete value
with a negative lookahead on a trailing digit.

* fix(aimlapi): make the one-shot key exchange crash-durable and per-operation

Three checkout-recovery correctness fixes:

- Persist the exchanged key under the CAS BEFORE returning it. The lease winner
  used to hand the /exchange key to the caller, which wrote the receipt only
  afterward; a crash in between left the checkout exchanged but its only key
  unpersisted, so a retry re-ran (and was rejected by) the spent one-shot
  exchange. exchangeKeyWithLease now records the settled receipt via
  recordAimlapiSettledKeyAsync (merges over the record, clears the lease) as soon
  as the exchange succeeds.

- Use a per-operation exchange-lease owner instead of a module-global id. Two
  overlapping top-ups in the same process shared one owner, which the acquire
  treats as self and immediately reclaims, so both could POST the non-idempotent
  /exchange concurrently. A fresh owner per operation makes the second observe
  the first's lease as foreign and back off; a retry within one operation keeps
  its owner and still reclaims the lease it released.

- Never serialize an empty apiKey/apiKeyId. The existing-key top-up path reports
  apiKeyId: '', which the reader rejects, making the whole settled receipt (and
  the paid key it records) unrecoverable. The save path now coerces an empty
  key/id to absent so the receipt stays readable.

* fix(aimlapi): refuse to overwrite an unfinished checkout when the intent changes

claimAimlapiTopupState backs a single slot, so rerunning with a different amount,
auto-top-up, or endpoint used to unconditionally replace the stored record. When
the prior checkout had opened a session (a resume token — possibly already paid
but not yet exchanged) or held a settled key not yet written to a profile, that
dropped the only handle to a paid session/key and stranded it permanently.

claim now refuses a changed intent while such a record exists, with an actionable
message to finish or cancel the earlier top-up first (re-running the same intent
still resumes it). A never-advanced claim — empty resume token, unsettled, no key
— is still replaced. The CLI surfaces the message directly; the interactive flow
already clears the prior record on edit, so normal re-edits are unaffected.

* fix(aimlapi): never settle a keyless receipt; keep the paid key reaching the profile

Addresses a further review batch:

- recordAimlapiSettledKeyAsync now refuses to mark a receipt settled (and clear
  the lease) when no key resolves from the call or the stored record. A keyless
  settled receipt would make a peer resume from a spent one-shot exchange with no
  credential; the record and its lease now survive so a retry can still exchange.

- The CLI's pre-profile settled-receipt save is now best-effort (try/catch + a dim
  note), matching the earlier saves. A lock/permission/IO failure there no longer
  throws before finishProfile, so the paid, exchanged key still reaches the
  provider profile.

- startCreateFromPreset drops aimlapiPersistedIntentRef on a fresh flow entry
  (in-memory only) so a later resetAimlapiCheckoutIntent can never clear a previous
  flow's on-disk receipt against a stale payment id.

- Prompt copy: "Do you have an aimlapi.com key?" / "I already have an aimlapi.com
  key" (missing article).

Tests: keyless settle is rejected and leaves the lease intact; the CLI forwards
explicit --amount/--model; the settled-receipt recovery renders the top-up (not
"ready") done copy.

* test(aimlapi): assert the exchange lease stays held on a keyless settle attempt

Tighten the keyless-settlement guard test: a "not settled" assertion also passes
if the lease were wrongly cleared (a peer would then see 'acquired'). Assert the
peer acquisition returns 'held' so the test pins that a keyless settle preserves
the lease for a retry.

* fix(aimlapi): poll the checkout token, not the auth bearer, while waiting on a resumed exchange

pollUntilExchangeSettled was called with the passwordless-auth bearer
instead of the partner checkout-session token, so it polled the wrong
resource. A terminal error there clears the recovery receipt, stranding
a paid one-shot sign-up exchange.

* fix(aimlapi): keep the checkout receipt resumable through an unconfirmed amount/auto-top-up edit

Editing the amount or auto-top-up cleared the persisted checkout intent
and durable receipt immediately, before the abandon-ack confirmation
that gates actually starting a new payment session. A user who edits
and backs out (or completes the still-open browser checkout) before
confirming lost the only mapping to that chargeable checkout, so a
later run would open a new one instead of resuming the paid session.

The reset now happens only once the user has explicitly confirmed
abandonment: claimAimlapiTopupState takes an `abandonExisting` option
that atomically overwrites the retained record under the same lock
acquisition, instead of racing a separate async clear against a
synchronous claim.

* test(aimlapi): sync on the rendered email before submitting in the new receipt-resume test

A fixed sleep doesn't guarantee the TextInput has processed the typed
email before Enter is sent; a loaded runner can drop the submit. Wait
for the typed value to actually render, matching the amount-edit sync
already used later in this same test.

* docs(aimlapi): describe checkout retention as durable, not session-scoped

The prior wording ("retained while the provider flow remains open")
undersold what topupState.ts actually does: the payment identity and
any issued key are persisted to disk, so a restart resumes the same
checkout too, and a prior paid+exchanged run finishes the profile
write on the next run instead of re-provisioning.

* fix(aimlapi): unify error-status extraction, drop dead top-up state, tighten wrappers

- Extract aimlapiApiErrorStatus as the one place that reads an HTTP
  status off a caught error, structurally (not `instanceof
  AimlapiApiError`) since some callers surface a duck-typed error with
  a bolted-on `status` instead of the real class; use it at both call
  sites that previously duplicated (and disagreed on) this check.
- Remove aimlapiPaymentSessionId/isAimlapiTopupRunning: both were
  write-only state (declared with a blank destructure slot, never
  read), so every setter call scheduled a render for no observable
  effect.
- Switch providerManagerAimlapi.ts's wrappers to `...args` forwarding
  so an implementation gaining a parameter can't silently get it
  dropped by a wrapper that still names the old ones positionally.
- Stop exporting pollUntilPaid from the aimlapi barrel; nothing
  imports it through there (topup.test.ts imports it directly from
  topup.js), so keep it out of the public surface.
- Normalize the email key while rebuilding the sign-in key store's
  collection branch on read, matching the legacy single-record
  migration branch right above it - a hand-edited or older-build file
  with a mixed-case key would otherwise be invisible to
  loadAimlapiSignInKey and mint a duplicate key.

* test(aimlapi): cover resetAimlapiCheckoutSession, by-key top-up args, and error edges

- resetAimlapiCheckoutSession: refreshes the payment session while
  preserving a minted key, and is a no-op when there's no key to
  preserve.
- ProviderManager: a low-balance saved key that gets topped up charges
  the EXISTING key via topUpAimlapiByApiKey (apiKey, non-empty
  paymentSessionId, empty resumeSessionToken) instead of opening a new
  passwordless-account checkout - previously only exercised through
  the default test mock, with no assertion on the call.
- The three negative assertions in the top-up progress-frame check
  tested strings that don't exist anywhere in this GUI (CLI-only or
  pure invention), so they could never fail; add a check against the
  real failure copy so a regression that silently fails at that point
  is actually caught.
- CLI: pin the --no-open default (false) when the flag is absent, and
  cover the non-Error (thrown string) branch of the handler's
  credential-redaction path - both previously only exercised through
  the Error/AimlapiApiError branches.

* fix(aimlapi): close checkout-state concurrency and exchange-lease races

- saveAimlapiTopupState now merges resumeSessionToken like the other
  retained fields instead of spreading the caller's value verbatim. A
  caller saves this record at points where its in-memory copy is still
  empty (right after sign-in, before a checkout session exists); a
  concurrent peer running the same intent can have already elected and
  recorded a real token in that window, and the unconditional spread
  was overwriting it with "", stranding the peer's chargeable checkout.
- The exchange lease is sized for a single POST (EXCHANGE_LEASE_STALE_MS,
  75s) but a resumed wait-exchange holder can sit in a read-only poll
  for up to POLL_TIMEOUT_MS (20 minutes) before ever reaching that POST.
  Without refreshing, a peer would see the lease go stale mid-wait,
  reclaim it, and risk a second concurrent /exchange on the same
  one-shot session. Add refreshAimlapiExchangeLeaseAsync and call it
  every poll iteration.
- When a peer finishes /exchange and records the settled key WHILE this
  process holds the lease and is polling/exchanging, the poll seeing
  the session flip to 'exchanged' threw a hard failure instead of
  resuming from that peer's settled receipt. Re-check for a settled
  receipt before releasing the lease and rethrowing.
- claimAimlapiTopupState's abandonExisting no longer drops an
  already-minted (but not yet paid) existing-account key when
  overwriting a retained checkout for a different amount/auto-top-up -
  it now merges apiKey/apiKeyId/model in, matching
  resetAimlapiCheckoutSession's retain-key pattern. A fully settled
  (paid + exchanged) credential is refused unconditionally regardless
  of abandonExisting, since that confirms giving up an UNPAID checkout,
  never an already-paid one.

* fix(aimlapi): guard GUI checkout abandonment and receipt recovery

- The abandon-ack gate only armed once a checkout URL surfaced
  (aimlapiOpenedCheckoutRef), but resolveTopupSession can already elect
  and persist a resumeSessionToken before that point. Backing out in
  that window then editing the amount hit claimAimlapiTopupState's
  generic refusal instead of the same confirm-to-abandon flow. Extend
  the gate to also cover a persisted (not yet opened) intent.
- Persist an existing-account key minted at sign-in into the top-up
  receipt itself (mirrors the CLI), not just the separate sign-in-key
  cache, so a restart before settlement can resume from one
  self-contained record instead of depending on two files staying
  consistent.
- reportSession's terminal branch (a cancelled/expired/dead session)
  always fully wiped the receipt; mirror the CLI's persistSession,
  which retains an already-minted key (fresh payment session, dead
  token dropped) and only falls back to a full clear when there's no
  key to keep.
- Submitting the email screen unconditionally reset the whole
  onboarding identity, silently abandoning a chargeable checkout on an
  accidental Esc-back-and-resubmit of the same email. Require the same
  explicit confirmation an amount edit does when a resumable checkout
  exists.
- "Set up a new key or switch account" only cleared in-memory fields,
  leaving a durable receipt from an earlier interrupted top-up (this
  mount's refs were never populated for it, since it may be from an
  earlier process) to hit the same refusal on the next onboarding
  attempt with no way to recover short of deleting the file by hand.
  Force the next claim to override it once.
- existingAimlapiCredential() rejected saved-profile discovery whenever
  the AMBIENT AIMLAPI_INFERENCE_URL wasn't canonical, even for a
  profile that was itself saved against the canonical endpoint. Narrow
  the canonical requirement to what it's actually protecting: reading
  the ambient env key, and sending a saved key to a non-canonical
  endpoint (the existing per-profile check).
- The post-signup success screen claimed a magic link was emailed; this
  flow is passwordless email-code sign-in, no magic link is ever sent.
  Point at the dashboard instead.

* docs(aimlapi): note the interactive/CLI auto-top-up default mismatch

The guided GUI flow pre-selects auto-top-up on; the CLI's --auto-top-up
only enrolls when explicitly passed. Left both defaults as-is (auto-top-up
is a real billing behavior, not something to flip unilaterally) and
documented the asymmetry so it's not a surprise either way.

* test(aimlapi): sync on the settled frame before confirming switch-account

A fixed sleep doesn't prove the Select's focus actually moved to the
second option; on a loaded runner the following Enter could land on
"Continue with your saved API key" instead and assert against the
wrong branch. Wait for the frame to stop changing, matching the
settle-poll pattern already used elsewhere in this file.

* fix(aimlapi): stop the exchange poll when a peer reclaims the lease

The periodic lease refresh added to pollUntilExchangeSettled discarded
its result (`.catch(() => false)`), so a peer reclaiming the lease
mid-wait was silently ignored: the poll kept going, returned normally
once the session left 'exchanging', and the caller walked straight
into the non-idempotent /exchange POST with no ownership check of its
own — racing whatever the peer was doing with the same one-shot
session. The comment claiming this was safe ("resolves on this
function's next outer retry") was simply wrong: there is no outer
retry on the success path, control goes directly to the POST.

Distinguish a thrown refresh (transient lock contention — best-effort,
retry next iteration) from an explicit `false` result (the lease is
definitively no longer ours) and bail out on the latter, so the
caller's existing catch block re-checks for the peer's settled
receipt (or fails the run, requiring a re-run) instead of racing it.

* test(aimlapi): require the settle-wait frame to actually differ from before the keypress

waitForCondition polls every 10ms; on a loaded runner two consecutive
polls can both land before Ink has processed the keypress at all, so
the "stable frame" check was satisfied by the unchanged PRE-keypress
frame, sending Enter before focus ever moved to the second option.
Snapshot the frame before the keypress and require the settled frame
to differ from it, not just be internally stable.

* fix(aimlapi): elect the retained key atomically, stop blocking Ink on claim

- Two concurrent sign-ins for the same intent could each mint their own
  existing-account key before either save landed, and saveAimlapiTopupState
  (last-writer-wins) let whichever saved last silently overwrite the
  other's key on disk while both runs kept using their own in-memory
  copy. Elect apiKey/apiKeyId first-writer-wins (same as
  resumeSessionToken already is), re-check the receipt right before
  minting so a losing run adopts the winner's key instead of minting a
  second, and re-check again after a save that lost the election so the
  run's own in-memory key matches what's actually on disk. Apply the
  same first-writer-wins election to the separate GUI sign-in-key cache
  (saveAimlapiSignInKey), which had the identical last-writer-wins gap.
- The GUI called the sync claimAimlapiTopupState directly from an event
  handler; its lock retry blocks the whole event loop (Atomics.wait) for
  up to LOCK_TIMEOUT_MS on contention, freezing Ink rendering, Esc, and
  SIGINT — exactly what resetAimlapiCheckoutSessionAsync already exists
  to avoid for the same reason. Add claimAimlapiTopupStateAsync (sharing
  the same claim logic via an extracted operation function) and switch
  the GUI to it, making startAimlapiTopup async.

recordAimlapiCheckoutSession (the reportSession/onSession path) has the
same sync-lock exposure but is called from a callback whose return value
AimlapiProvisionOptions.onSession drives synchronous control flow in
several places across both the CLI and GUI provisioning paths; making it
async is a larger, riskier contract change deliberately left out of this
pass.

* fix(aimlapi): never pair a new key with a stale or unrelated apiKeyId

saveAimlapiTopupState's apiKeyId fallback still read current.apiKeyId
even when current.apiKey was empty (the id-without-a-key case) or when
state carried a genuinely new apiKey with its own empty-id sentinel,
letting a fresh key get silently tagged with an unrelated leftover id.
Gate apiKeyId on the same winner apiKey came from instead of falling
back to current independently.

* fix(aimlapi): stop cross-account key leaks, lease key-minting, keep GUI CAS async

- claimAimlapiTopupState's abandonExisting carried a retained apiKey into
  ANY differing intent, including a switch from account A to account B
  (the GUI's forceAbandonExisting path). A B-flow restart before the
  profile write could then initialize from the receipt and call the B
  checkout with A's credential — crediting A while B's flow saves A's
  key. Gate the carry-over on the intent's account/key identity
  (`email`) matching, not just abandonExisting.
- The key-choice screen (I am a new user / I already have a key) reset
  the whole onboarding identity unconditionally on either choice, even
  when Esc had backed all the way out from the amount screen past an
  already-opened, still-chargeable checkout. Apply the same
  abandon-confirmation gate startAimlapiEmailOnboarding already uses.
- POST /v1/keys (minting an existing-account key) had no cross-process
  serialization: two concurrent runs for the same intent could each
  observe no retained key and both mint, orphaning whichever key lost
  the first-writer-wins receipt race. Add a key-mint lease (mirroring
  the exchange lease's acquire/release shape) so exactly one process
  ever mints; a peer backs off and adopts the winner's recorded key.
- The interactive flow already claimed asynchronously, but still called
  the synchronous saveAimlapiTopupState and recordAimlapiCheckoutSession
  directly from an event handler and the onSession callback — either
  could block the whole event loop for up to LOCK_TIMEOUT_MS under lock
  contention, freezing rendering, Esc, and SIGINT while a payment
  session is being created. Add async CAS variants and await them; this
  needed widening AimlapiProvisionOptions.onSession to allow returning a
  promise, since its return value drives resolveTopupSession's session
  election.
- An ambient AIMLAPI_API_KEY takes the by-key route with
  aimlapiExistingUsesEnv, so the eventual profile intentionally stays
  keyless. The settled-receipt save before that write unconditionally
  copied the env value into aimlapi-topup.json regardless, expanding a
  secret's on-disk exposure surface for no recovery benefit (a restart
  re-reads the same env var). Keep an env-backed receipt credential-free.

* fix(aimlapi): preserve the key-mint lease across unrelated CAS writes

saveTopupStateOperation and recordCheckoutSessionOperation merged the
exchange lease but not the key-mint lease added in the previous commit:
AimlapiCheckoutState (what every caller spreads checkoutState from)
carries neither lease pair, so an unrelated write - persisting the
exchange flag, or electing a checkout session - silently dropped an
in-flight peer's key-mint lease. A third process would then see the
slot as free and mint its own key, reopening the exact double-mint race
the lease exists to close. Fall back to the current lease the same way
the exchange lease already does.

Also: add a future-dated key-mint lease reclaim test mirroring the
exchange lease's, and align the default saveAimlapiTopupStateAsync test
mock with the real CAS (match on intent + payment id, keep the first
writer's resumeSessionToken/apiKey) so it no longer accepts a write the
real store would reject.

* test(aimlapi): preserve the key-mint/exchange lease in the mocked GUI CAS writes

saveAimlapiTopupStateAsync's and recordAimlapiCheckoutSessionAsync's
default mocks spread { ...state } as their write's base, same gap as
the real saveTopupStateOperation/recordCheckoutSessionOperation had
before the previous commit: neither lease pair survived a write whose
state didn't carry them (which is every real caller, since
AimlapiCheckoutState exposes neither).

Fixing the merge alone wasn't enough — the same two mocks' "does this
write still belong to this slot" check also compared lease fields as
if they were part of the intent identity, so a write that seeded a
lease value failed to match the just-claimed record and silently
no-op'd instead of persisting anything. Exclude both lease pairs from
that comparison too, matching the real matchingStateOrNull (which only
ever compares INTENT_KEYS + paymentSessionId).

* fix(aimlapi): recover ambiguous key-mint/exchange outcomes before releasing leases

createKey and /exchange are both non-idempotent with no server-side retrieval
path, so a lost response after the request actually committed left three
races: a retry could exchange (or mint) a second time and orphan the first
credential, or the CLI's exchange caller would surface a generic network
error instead of the accurate "already exchanged, rotate the key" guidance.

exchangeKeyWithLease now distinguishes a genuinely ambiguous transport
failure of the /exchange POST itself from other doExchange failures (a
pre-POST bail on a reclaimed lease, or a definite rejection): only the
former re-checks the session status directly, surfaces the already-exchanged
error when confirmed, and otherwise leaves the lease held instead of
releasing it into a race. mintExistingAccountKeyWithLease applies the same
ambiguous/definite split before deciding whether to release its lease.

The GUI sign-in flow had an equivalent gap one step earlier: two concurrent
code-verification races could each see an empty key cache and both mint
before either save elected a winner, so the loser never adopted the winner's
key. completeAimlapiCodeSignIn now serializes the cache lookup and mint
behind a new email-scoped lease in topupState.ts, so a losing process waits
and adopts the winner's cached credential instead of minting its own.

* fix(aimlapi): treat caller-aborted mutations as ambiguous and dedupe the transport helpers

client.request rethrows a caller-driven abort as the raw abort error instead
of wrapping it in AimlapiApiError, so the ambiguous-outcome checks added for
createKey and /exchange missed it: cancelling client-side does not stop a
non-idempotent POST from completing server-side, but the lease was still
released as if the request definitely failed, leaving the door open to a
retry racing a second mint/exchange. All three call sites (the checkout-time
key-mint lease, the exchange lease, and the sign-in key-mint lease) now also
hold the lease when the caller's own signal fired.

Extracted the duplicated abortError/sleep/isAmbiguousTransportApiError
helpers shared between topup.ts and onboarding.ts into transport.ts so the
ambiguity rule can't drift between the CLI and GUI paths. Switched the GUI
sign-in flow's cache save to the async, lock-yielding variant and logged its
lease-release failures for parity with the other leases.

* fix(aimlapi): close six checkout-state races found across the claim, lease, and recovery paths

claimAimlapiTopupState's in-progress check only looked at
resumeSessionToken/settled/apiKey, so a receipt claimed just before its
non-idempotent POST (/v1/keys or /exchange) still looked blank and
replaceable to a different intent. A competing claim could overwrite it
mid-flight, leaving the in-flight request's eventual CAS save with no
matching record to land in and orphaning the credential it was about to
mint or exchange. The claim now also refuses (unconditionally, even under
abandonExisting) while either lease is live.

The sign-in key-mint lease's 75s stale window exactly matched createKey's
worst-case duration (60s) plus the async lock's own timeout (15s) for the
cache write that follows, with zero margin for anything else. A legitimately
still-working holder could lose the lease to a peer moments before its
result was cached. It's now refreshed right after createKey succeeds, giving
the cache-write phase its own fresh window.

ProviderManager's code-verification path called the synchronous
saveAimlapiSignInKey, whose lock retry blocks the whole event loop for up to
five seconds on contention — freezing Ink rendering, timers, Esc, and SIGINT
right after a sign-in. Switched to the async variant, exported through
providerManagerAimlapi.ts alongside the other async cache operations.

The three session polling helpers typed onSession as returning void and
never awaited it, even though ProviderManager's callback is async and starts
receipt cleanup before returning. A terminal session (cancelled/expired/
failed, or a dead session) could let the UI reach the amount screen before
the durable receipt was actually reset, so an immediate retry still saw the
stale resume token and got rejected as "not yet abandoned." Both sides now
await through to completion.

The confirmed email-switch flow cleared the in-memory checkout intent and
fired an un-awaited, error-swallowing state clear, but derived its later
claim's abandonExisting only from refs that clear had just wiped — so a
slow or failed clear left the user's explicit confirmation unenforced at the
claim itself. It now sets the same one-shot force-abandon signal the
"switch account" flow already uses for exactly this kind of on-disk,
this-mount-invisible conflict.

A cached sign-in key that the server had revoked was indistinguishable from
one that was merely unreachable: both collapsed into balanceStatus:
'unknown', which re-cached the same dead key and sent the user to manual-key
entry with no way back into the guided flow short of deleting local state.
A definite 401/403 against a cached (not freshly minted) key now invalidates
the stale cache entry and mints one replacement before falling back to the
generic unknown-balance path; every other (ambiguous) failure still leaves
the cache untouched.

Extracted a shared claim/lease-liveness helper in topupState.ts and added
regression coverage for each race — including two that hold a mocked
createKey/reset call open to prove the competing operation actually waits
instead of just asserting on the end state.

* fix(aimlapi): clear the stale force-abandon flag on a fresh preset entry

aimlapiForceAbandonExistingRef is armed when the user confirms abandoning a
checkout during an email switch, then consumed by the next claim. If that
claim never runs — the switch's own onboarding fails and the user backs all
the way out to preset selection instead of retrying — the flag stayed armed.
Re-entering the aimlapi preset with an unrelated email then passed
abandonExisting: true on its first claim with no confirmation for that flow,
silently overwriting whatever unpaid checkout was still on disk.
startCreateFromPreset now resets the flag alongside the other per-flow refs
it already clears on fresh entry.

Also swapped a fixed 20ms sleep in the cross-intent concurrency test for a
signal fired from the held-open /v1/keys handler, so the test can't flake
under CI load waiting for the run to reach the point it needs to race.

* fix(aimlapi): close the remaining confirmation, cleanup, and lease gaps in checkout state

The API-key-choice screen's own confirm-abandon gate (Enter twice to accept
"a checkout from this account is still pending") reset the onboarding
identity but never armed the force-abandon signal the email-switch and
switch-account flows already use. A contended or failed pre-clear left the
next claim to hit the CAS's unconfirmed-conflict refusal despite the user
having just confirmed abandonment through this exact screen.

reportSession('')'s terminal-session handler discarded the persisted intent
ref before its reset/clear attempt settled, and swallowed any failure as
success. A lock timeout or I/O error then left the durable receipt exactly
as it was, but with no ownership left in memory to retry cleanup or to route
a later conflicting claim through the normal confirmation gate — the CAS
just rejected it outright. Ownership now only drops once the transition
actually commits; a failure is logged and the ref stays populated so the
existing gate covers the next claim.

The sign-in key-mint lease's stale window already had zero margin for its
own refresh call's lock wait (up to 15s) on top of createKey's own worst
case (60s) and the cache save's lock wait (another 15s) — 90s with nothing
left over. Widened it to 150s and lengthened the losing side's patience to
match, and stopped silently ignoring a refresh that reports lost ownership:
it's now logged for diagnosability even though the save itself stays safe
to attempt regardless (first-writer-wins makes a losing write a no-op).

Both onSaved completion paths (persistExistingAimlapi and persistAimlapiKey)
called the synchronous clearAimlapiSignInKey from Ink's synchronous save
callback, whose lock retry blocks the event loop for up to five seconds on
contention — freezing rendering, timers, Esc, and SIGINT right at
completion, the same class of bug already fixed for the sign-in save path.
Re-exported the async variant through providerManagerAimlapi.ts and switched
both call sites to fire-and-forget it instead.

* fix(aimlapi): stop the flow instead of risking a stranded key on a receipt-write failure

/exchange (and the by-key top-up) is a one-shot operation: once it succeeds,
the issued key exists only in memory until a durable copy lands somewhere.
Both the CLI and the GUI wrote the local recovery receipt right after that,
but treated a failure there as best-effort and proceeded straight into the
provider-profile write regardless. If the receipt write failed and the
profile write then also failed — or the process was interrupted between the
two — the key was gone: nothing durable ever recorded it, and a retry can't
re-exchange an already-spent session to get it back.

Both paths now treat the receipt as a required checkpoint rather than an
optional resume aid: a failure here stops the flow with a clear error
pointing at the one real recovery path (rotating the key from the aimlapi.com
dashboard) instead of silently continuing. This shouldn't cost much in
practice — the underlying CAS write already retries substantially on lock
contention before giving up, so a failure this deep signals a real problem
rather than a transient blip a fallback write would likely have hit too.

Added failure-injection coverage for both paths: the CLI test breaks the
config directory (a file where a directory is expected) right as the
exchange response lands, so the post-exchange save fails deterministically
without relying on OS-specific permission semantics; the GUI test mocks the
save to reject directly and asserts the profile write is never reached.

* test(aimlapi): strengthen receipt-write-failure coverage and document the recovery path

Both the CLI and interactive-flow tests for the post-exchange receipt-write
failure only asserted the generic error text, which would still pass if the
issued key id — the actual recovery handle the error exists to surface — got
dropped from the message later. Both now assert the id appears too. The
interactive test also confirms the screen stays usable after the error: a
retry reaches the amount-submission path again instead of the flow being
stuck.

Documented the resulting behavior in the setup guide: since the key exchange
is one-shot, a receipt-write failure after a successful payment now stops
both flows with an error naming the issued key, rather than continuing
silently — recovery is manual, via rotating that key on the aimlapi.com
dashboard.

* test(aimlapi): move to end-of-line before clearing the email field after Esc

Rebasing onto current main picked up the input layer's DEL-coalescing fix,
which now correctly respects the cursor position for a backspace run instead
of dropping it. These three tests backspaced assuming the cursor sat at the
end of the retained email text, but cursorOffset is a single state shared
across every screen's text field and was last set for the amount screen (its
default "25" is 2 chars) — going back via Esc never resets it, so the cursor
was actually stuck mid-string. Sending an explicit end-of-line sequence
before the backspaces makes the clear correct regardless of where the stale
cursor was left.

* fix(aimlapi): close six checkout/onboarding gaps from the latest review pass

Treats a malformed-but-2xx key-mint response as ambiguous (not proof of
failure) so an unusable receipt no longer releases the mint lease and risks
an orphaned credential; fences startAimlapiTopup's cancellation to an
epoch created before the state-lock await so Esc/unmount during that wait
can no longer barge back in; makes the checkout-receipt read fail closed on
a permission/IO/parse/schema failure instead of silently claiming over it;
completes (or reconciles) a settled by-key receipt instead of stranding it
at the post-payment model picker; retires a sign-in mint lease together
with its cache entry so it can't resurface as held once the cache is later
cleared; and adds a --code-stdin path plus a deprecation warning so the
passwordless code no longer has to travel through shell history or argv.

* fix(aimlapi): reconcile env-credential receipts and stop endorsing AIMLAPI_CODE as safe

reconcileSettledAimlapiTopupStateAsync matched a stale settled receipt by
its stored apiKey, but an env-sourced credential's receipt never persists
one, so that path stayed permanently stranded; it now matches on the
absence of a stored key when the caller is reusing an env credential.
Also stops recommending AIMLAPI_CODE as an equivalently safe alternative
to the deprecated --code flag, since typing it inline still lands in
shell history, and adds a lock-release assertion to the fail-closed
receipt-read tests.

* fix(aimlapi): await stale-receipt reconciliation and stop overclaiming --code-stdin's history safety

Reconciling a stale settled receipt matches by the by-key credential's
apiKey, which a genuinely new top-up for that same credential can also
produce — firing the reconcile call without waiting for it left a window
where a fresh settlement landing in that gap could be swept up by it.
Await it before moving on so the two can no longer interleave.

Also narrows the --code-stdin messaging: it only guarantees the code stays
out of this process's argv/`ps` output, not shell history in general,
since that still depends on how the caller feeds stdin.

* fix(aimlapi): keep the configured screen locked until reconciliation finishes

Clearing isAimlapiKeyValidating before the reconcile await let the
aimlapi-configured screen's Select (and its Esc binding) become
interactive while that reconcile was still running in the background —
it carries no abort signal of its own, so aborting the surrounding
controller only stops this flow from acting on the result, not the
reconcile itself. That left a window where the user could start a
competing top-up for the same credential and have its fresh settlement
caught by the still-in-flight reconcile. Both now stay gated on
isAimlapiKeyValidating through the whole wait.

* fix(aimlapi): fail closed on the sign-in cache, bind env receipts by identity, and commit minted keys

Makes the sign-in key cache and its mint lease match the checkout
receipt's fail-closed contract: only ENOENT means no record, so a
permission/IO/parse failure can no longer be mistaken for "nothing
cached, no lease held" and authorize a second createKey call or a
concurrent lease acquisition.

Reworks reconcileSettledAimlapiTopupStateAsync to match on the by-key
checkout intent's non-secret key fingerprint (already carried in the
persisted email field) instead of the raw stored apiKey or its mere
absence — an env-backed receipt never stores its key, so absence alone
couldn't tell two different env credentials' receipts apart, letting one
credential's balance check discard another's still-unrecovered payment.

Treats persisting a freshly minted existing-account key as a commit
point in the CLI's mint-with-lease path: a write failure now stops the
flow with a recovery-oriented error and leaves the lease held, instead
of continuing with the key only in memory where an interruption before
the later checkout/profile save would orphan it once the lease goes
stale.

Adds a shared isValidAimlapiSignInCode check so both the CLI and the
interactive flow reject a malformed passwordless code (empty,
non-numeric, wrong length) before it ever reaches verifySignInCode.

* fix(aimlapi): reject an array-shaped sign-in cache/lease file instead of degrading to empty

typeof [] === 'object' and [] !== null, so readJsonObjectFile's shape
check let a JSON array through as if it were a valid store. Both readers
then found no matching entries and returned {}, exactly the "no cached
key, no live lease" outcome the fail-closed contract exists to prevent —
authorizing a second createKey call or a lease acquisition over a
possibly-live one.

* fix(aimlapi): commit the sign-in key as a checkpoint and retire a completed mint's lease

mintOrAdoptSignInKey swallowed a failed cache commit and returned the
minted key only in memory — the same non-idempotent-mutation gap already
closed for the CLI's mintExistingAccountKeyWithLease. A commit failure
now stops the flow with a recovery-oriented error and leaves the lease
held, instead of risking the key becoming unrecoverable once the lease
ages out and a retry mints a second one.

Also retires the checkout key-mint lease in the same save that elects a
freshly minted key, mirroring how the exchange lease is already cleared
on settle. Without it, backing out of an unpaid checkout and confirming
a different amount right away still hit the "minting or exchanging"
refusal for the full 75s stale window even though the mint had already
completed.

* fix(aimlapi): make key-mint lease retirement owner-checked, preserve error causes

Retiring the checkout key-mint lease on a successful mint (the previous
commit's fix) cleared it unconditionally, with no check that the save
still belonged to the owner that acquired it. createKey has no refresh
mechanism, so a slow response can let the lease go stale and be
reclaimed by a peer before the original owner's save lands — clearing
the lease then would drop that peer's still-live one and let a
differently-amounted claim proceed as though minting were done while
the peer's mint was still genuinely in flight.

Replaces the raw saveAimlapiTopupState call in
mintExistingAccountKeyWithLease with a dedicated
recordAimlapiMintedKeyAsync that takes the acquiring owner and only
retires the lease while it's still theirs — mirroring how
recordAimlapiSettledKeyAsync already handles the exchange lease.

Also adds `cause` to the three recovery-oriented errors thrown on a
receipt-write failure, so the underlying persistence error stays
diagnosable instead of being replaced by the wrapper message alone.

* test(aimlapi): assert a stale owner's minted key is still persisted

The reclaimed-peer-lease regression test only pinned the lease
bookkeeping, so a variant regression that skipped the write entirely for
a non-owning caller (e.g. an early return on lease-owner mismatch) would
still pass while silently discarding a real, non-idempotently minted
key. Asserts the receipt retains it regardless of who currently holds
the lease.

* fix(aimlapi): persist and charge the endpoint a manually-entered key was actually validated against

persistExistingAimlapi's no-existing-profile fallback saved the profile
with resolveEndpoints().inferenceBaseUrl (the current ambient endpoint)
instead of aimlapiInferenceBaseUrl (the endpoint this flow actually
validated the key and will charge against). The two diverge for a
manually-entered key: after "Set up a new key or switch account" resets
aimlapiInferenceBaseUrl to the ambient default, draft.baseUrl (what
validateAndPersistAimlapiKey actually calls the balance/top-up endpoints
with) keeps the OLD profile's endpoint if it differs (e.g. a canonical
saved profile while AIMLAPI_INFERENCE_URL currently points at a proxy).

Now captures the validated endpoint into aimlapiInferenceBaseUrl as soon
as the low-balance branch is reached, and the fallback save uses that
state instead of re-resolving the ambient endpoint — so the top-up
charge and the persisted profile both follow the endpoint that was
actually validated.

* fix(aimlapi): reject stale key-mint results, make the exchange checkpoint mandatory, validate lease pairs

recordAimlapiMintedKeyAsync previously let a stale owner's delayed
createKey result land beside a peer's reclaimed, still-live lease: since
no key was recorded yet, first-writer-wins accepted the stale result
outright, so the reclaiming peer's own (equally real, non-idempotent)
mint got silently discarded once its own save landed — turning one lost
credential into two. It now rejects a result whose ownership was already
lost when nothing is recorded yet to adopt instead, surfacing a
recovery-oriented error naming the issued key id rather than risking a
second orphan. The caller now also returns whichever credential is
actually durably recorded, not always its own.

exchangeKeyWithLease's own settled-receipt commit — the only durable
record of a one-shot /exchange result until the caller's later,
separate save runs — was only logged on failure. A crash in that
window left the paid session exchanged with its key absent from local
recovery state, so a retry could only report the session was already
exchanged with no way to recover automatically. The commit is now a
required checkpoint: its failure stops the flow with the existing
recovery guidance and leaves the exchange lease held, exactly as the
analogous key-mint checkpoint already does.

The receipt schema validated the exchange lease pair but not the newer
key-mint one, and even the exchange check only verified each field in
isolation — a one-sided pair (an owner with no timestamp, or vice versa)
passed either way. A shared validator now enforces both lease pairs are
either fully present or fully absent, so a malformed or partial lease
can no longer be silently accepted as "not currently live" and let a
claim replace it out from under an in-flight mint or exchange.

* fix(aimlapi): fail the exchange when the settled-receipt commit no-ops, not just when it throws

recordAimlapiSettledKeyAsync silently returned without writing whenever the
CAS no longer matched (checkout cleared/reset mid-flight) or no credential
could be resolved to settle with. exchangeKeyWithLease only caught thrown
errors, so both no-op paths let a successfully exchanged key return with no
durable local receipt. The function now returns a boolean, and the caller
treats false exactly like a thrown save error.

Also fixes recordAimlapiMintedKeyAsync returning the raw untrimmed apiKeyId
instead of the trimmed value it actually persisted.

---------

Co-authored-by: Lookoff123 <bataryshkinairina@gmail.com>
2026-08-14 10:12:13 +08:00
575b407275 feat(partners): add ApiSmart, refresh Novita AI logo (#2121)
- Add ApiSmart (https://www.apismart.ai) to the README partners table
  and the web partner strip, with a dark-theme logo variant (near-black
  wordmark recolored to white, white matte removed).
- Replace the Novita AI PNG logo with the new SVG wordmark plus a
  generated dark variant, wired through the same prefers-color-scheme
  <picture> pattern (README) and logoDark field (web).

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-08-13 11:11:43 +08:00
JATMNandGitHub 7cae4089d6 feat(xai): add Grok 4.6/4.5 to catalog, xAI provider, and gateways (#2117)
* feat(xai): add Grok 4.6/4.5 and hybrid catalog discovery

xAI's current flagship is Grok 4.6; keep shared capability flags so gateways can reference the new models, and let /v1/models surface later Grok IDs without another catalog bump.

* fix(xai): authenticate hybrid discovery for OAuth and keep gateway aliases aligned

OAuth xAI sessions had no /v1/models credential, so hybrid refresh failed; also drop curated alias IDs from discovery and map grok-build-latest on Atlas/Hicap to Grok 4.5.

* fix(xai): align OAuth hybrid discovery cache with runtime metadata

OAuth-only xAI sessions hashed the access token into discovery writes, but
runtime limit lookups only used env credentials, so uncataloged Grok IDs
fell back to default windows. Mirror stored OAuth on cache reads and isolate
discovery tests from the shared config home.

* fix(xai): keep Grok 4.6 PR on catalog and hybrid discovery

Move OAuth /v1/models auth and cache-key alignment out of this branch so the catalog, xAI hybrid vendor, and gateway references stay reviewable on their own.

* fix(xai): read discovered model context length

* fix(xai): preserve OAuth discovery metadata

* fix(xai): tolerate malformed discovery models

* fix(discovery): harden credential cache partitions

* fix(xai): unify discovery OAuth cache handling

* fix(xai): preserve OAuth discovery cache identity

* fix(discovery): use native cache namespace hashing

* fix(xai): keep OAuth discovery nonblocking

* fix(discovery): use opaque cache fingerprints
2026-08-13 10:43:41 +08:00
OpenClaude 6277bfadbc fix: Ling 3.0 Tiny :free window back to Aug 13 (official promo end) 2026-08-12 11:18:24 +08:00
OpenClaude ee64d80c2e fix: extend Ling 3.0 Tiny :free availability window to Aug 17 2026-08-12 09:06:43 +08:00
JATMNandGitHub c40d663b70 fix(web): link release notes to GitHub (#2114)
* fix(web): link release notes to GitHub

* test(web): verify release link safety contract
2026-08-12 08:26:03 +08:00
github-actions[bot]GitHubgithub-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
6e30b40de0 chore(main): release 0.28.0 (#2090)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
v0.28.0
2026-08-11 21:16:25 +08:00
JATMNandGitHub 5f8e7d101b Revert "fix(release): synchronize web changelog entries (#2088)" (#2113)
This reverts commit 7743cf280e.
2026-08-11 20:23:26 +08:00
7b03ad19a4 feat(opengateway): add Ling 3.0 Tiny :free — Day-0 launch, free until August 13 (#2112)
* feat(opengateway): add Ling 3.0 Tiny :free — Day-0 launch, free until Aug 13

inclusionai/ling-3.0-tiny:free (7.9B MoE, ~1.3B active, 262k ctx)
joins the picker via the gateway's OpenRouter wiring. The gateway
time-boxes it (free through Aug 13, rate limited) and delists it
server-side when the window closes.

* test(opengateway): Ling Tiny gateway mapping test + explicit window note

Addresses CodeRabbit review on #2112:
- new ling-tiny.test.ts (macaron.test.ts pattern) asserting the
  opengateway-ling-3.0-tiny-free entry maps both apiName and
  modelDescriptorId to inclusionai/ling-3.0-tiny:free, plus descriptor
  capabilities and runtime limits
- catalog note now dated explicitly ('Free through August 13, 2026')
  and the entry's lifecycle documented: the gateway time-boxes the id
  server-side and 400s after the window; this static catalog has no
  expiry mechanism (Ling Flash precedent), so the entry is removed or
  updated at window close

* feat(integrations): availableUntil catalog-entry expiry + Tiny lifecycle guard

Addresses CodeRabbit round 2 on #2112:
- new optional ModelCatalogEntry.availableUntil (ISO-8601): entries past
  the cutoff are dropped in getCatalogEntriesForRoute, the single choke
  point behind the model picker, gateway catalogs, and runtime limits;
  the pre-existing (previously unenforced) hidden flag is honored in the
  same filter; unparseable dates fail open
- the Ling Tiny entry sets availableUntil to the gateway's window end
  (2026-08-13T10:00:00Z), so the picker drops it the instant the
  gateway starts rejecting the id — no client release needed
- boundary regression test on both sides of the cutoff, and the picker
  expected-list test pins the clock inside the window (setSystemTime)
  so it stays deterministic after the date passes
- ling-tiny.test.ts now also asserts supportsPreciseTokenCount: false

* test(integrations): exact-cutoff + hidden + malformed-date coverage; fix stale lifecycle comment

Addresses CodeRabbit round 3 on #2112:
- ling-tiny.test.ts asserts the boundary at exactly 2026-08-13T10:00:00Z
  (cutoff is exclusive: entry already gone at that instant)
- registry.test.ts covers the two previously-untested filter branches:
  hidden entries dropped, availableUntil expiry (before / at / after
  cutoff), and a malformed availableUntil failing open
- the catalog comment above the Ling Tiny entry no longer claims the
  static catalog has no expiry mechanism — availableUntil is the guard

Validation commands run locally:
  bun run integrations:generate
  bun test src/integrations src/commands/model/model.test.tsx src/utils/model
  bunx tsc --noEmit
558 tests pass, typecheck clean.

* fix(model): route static picker entries through the availability filter

Addresses jatmn's P1 on #2112: model.tsx read catalog.models directly,
bypassing the availableUntil/hidden filter that only lived in
getCatalogEntriesForRoute — so after 2026-08-13T10:00Z the /model
picker would still offer inclusionai/ling-3.0-tiny:free and selecting
it would persist an id the gateway 400s.

- registry.ts exports filterAvailableCatalogEntries (shared with
  getCatalogEntriesForRoute)
- model.tsx filters the static entries AND the static+discovery merged
  list, so discovery-sourced entries with their own markers are covered
- routeMetadata.ts getRouteDefaultModel's catalog fallback filters too,
  so an expired entry can never become the implicit default
- new picker regression test pinned just past the cutoff asserts the
  expired entry is gone while the rest of the catalog is untouched

Validation: bun run integrations:generate; bun test src/integrations
src/commands/model/model.test.tsx src/utils/model (559 pass); bunx tsc
--noEmit (clean).

* fix(model): merge raw static entries so expired ones mask cached duplicates

Addresses CodeRabbit round 4 on #2112: filtering static entries before
mergeRouteCatalogEntries let a cached discovery entry with the same
apiName (and no availableUntil marker) re-enter the merged list, where
the post-merge filter could not remove it. The merge now takes the RAW
static list — the expired static entry wins the apiName dedup and the
post-merge filter then drops it, so neither copy survives. The filtered
list still drives the non-discovery path.

Regression tests in routeCatalogOptions.test.ts cover the cached
duplicate after the cutoff (including documenting the buggy pre-filter
order) and the masking inside the window.

Validation: bun run integrations:generate; bun test src/integrations
src/commands/model/model.test.tsx src/utils/model (561 pass); bunx tsc
--noEmit (clean).

* test(integrations): default-model fallback skips hidden and expired entries

Addresses CodeRabbit round 5 on #2112: getRouteDefaultModel's catalog
fallback changed in the availability-filter fix but had no focused
coverage. New routeMetadata.test.ts case (self-contained registry
mutation with the shared lock, mirroring registry.test.ts) verifies a
hidden default-marked entry and a past-cutoff availableUntil entry are
both skipped in favor of the remaining valid entry, and that a catalog
with nothing valid yields undefined rather than a rejected id.

Validation: bun test ./src/integrations/routeMetadata.test.ts (63
pass); bun run integrations:generate; bun test src/integrations
src/commands/model/model.test.tsx src/utils/model (562 pass); bunx tsc
--noEmit (clean).

* test(integrations): release shared mutation lock even if registry restore throws

Addresses CodeRabbit round 6 on #2112: the fallback test's finally
block ran _clearRegistryForTesting/ensureIntegrationsLoaded before
releaseSharedMutationLock, so a throw there would leave the lock held
and block later tests. Nested try/finally, matching registry.test.ts's
afterEach shape.

Validation: bun test ./src/integrations/routeMetadata.test.ts (63
pass); full related suites 562 pass; tsc clean.

---------

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-08-11 20:18:40 +08:00
BogdanandGitHub 16e332e108 fix(query): use monotonic watchdog deadlines (#2110)
Wall-clock corrections can otherwise manufacture an immediate timeout or postpone an already-scheduled one. Keep lifecycle timestamps in wall time while measuring deadline state and elapsed duration from a monotonic clock.\n\nRefs #1830
2026-08-11 19:09:26 +08:00
JATMNandGitHub ff91642364 feat(integrations): add ApiSmart OpenAI-compatible gateway provider (#2109)
* feat(integrations): add ApiSmart OpenAI-compatible gateway provider

Add a hybrid-catalog ApiSmart gateway with dedicated APISMART_API_KEY/APISMART_MODEL
env wiring, route detection, profile persistence, and regression tests so the
provider works via --provider, env-only setup, and saved profiles.

* fix(apismart): protect dedicated credentials

* fix(apismart): enforce route credential boundaries

* fix(apismart): honor explicit competing route models

* fix(apismart): validate credentials and default profiles

* test(apismart): isolate env-only client tests

* fix(apismart): align routing and validation contracts

* fix(apismart): centralize route capability contracts

* Revert "fix(apismart): centralize route capability contracts"

This reverts commit f8cd975e68.

* fix(apismart): reject template credential placeholders at the shared root

Expand the shared credential usability helpers so dotenv sentinels like
null/undefined cannot win ApiSmart env-only precedence or get mirrored into
OPENAI_API_KEY, and restore focused ApiSmart docs after dropping the
shared routing centralization.

* fix(apismart): reject null base-URL and placeholder profile keys

Treat dotenv `null`/`undefined` OPENAI_BASE_URL sentinels as unset so
--provider apismart can apply defaults, and align profile/validation
credential checks with the shared placeholder contract.

* fix(apismart): match AIMLAPI env-only intent and proxy credential withholding

Use the OpenAI-compatible env-only gate so lingering CLAUDE_CODE_USE_OPENAI still keeps ApiSmart identity, retain route id on retargeted profiles, and withhold ambient credentials on non-canonical relaunches.

* fix(apismart): restore AIMLAPI credential and canonical URL parity

Backfill APISMART_API_KEY on relaunch and keyless canonical profiles, and gate credential forwarding on an exact /v1 inference URL so dedicatedCredentialsOnly auth and ambient keys stay aligned with AIMLAPI.

* fix: simplify apismart gateway integration

* fix: reject credential placeholders consistently

* test: remove obsolete apismart exception coverage

* test: remove stale apismart model fixture

* fix(apismart): restore dedicated provider contract

* fix(apismart): enforce credential boundaries

* fix(apismart): clear inherited auth headers

* test(apismart): strengthen credential boundaries
2026-08-11 19:08:38 +08:00
JATMNandGitHub 7743cf280e fix(release): synchronize web changelog entries (#2088)
* fix(release): sync web changelog entries from release please

* fix(release): provide sync push credentials

* fix(release): gate and finalize web release sync

* fix(release): make web release sync recoverable

* fix(release): target release PR commands explicitly

* fix(release): honor manifest release configuration

* fix(release): recover failed web sync retries

* ci(release): run full sync preflight

* fix(release): validate synchronized PR head

* fix(release): validate bot sync in release job

* fix(release): harden bot-owned web sync

* fix(release): validate exact bot PR head

* fix(release): isolate and bind PR synchronization

* fix(release): isolate validation from write credentials

* fix(release): reject non-regular generated inputs

* fix(release): require a valid forward version bump

* fix(release): scope sync artifacts to run attempts

* fix(release): reuse validated artifacts across retries

* fix(release): resume readiness after a completed push

* fix(release): bind retries to the validated commit

* fix(release): address synchronization review findings

* fix(release): support CRLF changelog recovery

* fix(release): keep validation transitions fail-closed

* fix(release): restore scoped web sync and marker ownership

Cut the multi-job finalize state machine back to a single draft-until-push
sync path, and fix consecutive releases leaving stacked automation markers
by stripping leftover draft ownership when inserting the next version.

* fix(release): keep web sync from blocking npm publish

Move pending Release Please web sync into its own job so a sync failure cannot skip install-verify, npm, or docker after a release tag is already created.

* fix(release): close web-sync trust and policy gaps

Remove hand-curation escape hatches, split read-only validation from
write-only push, discover bot PRs by branch identity, and run the full
local gate suite before marking the release PR ready.

* fix(release): validate gates against synchronized commit

Commit the synced releases.ts in the read-only validate job before
typecheck, security scan, and whitespace checks so those gates inspect
the content that will be marked ready, not the pre-sync HEAD.

* fix(release): harden web-sync trust boundary and draft gating

Run sync from trusted main with only changelog/manifest overlaid from
the bot PR, re-draft after release-please, serialize sync without
canceling in-flight pushes, and require an explicit sync base.

* fix(release): restore overlaid inputs before validate cleanliness gate

Fetching changelog/manifest from the bot PR dirtied tracked files on the
trusted main checkout and made the final git-diff gate fail on every
pending release. Restore those overlays after sync and fetch origin/main
for the security/whitespace checks.

* fix(release): reuse validated sync artifacts on retry

* fix(release): validate release sync inputs and retries

* fix(release): recover web sync state transitions

* fix(release): protect generated release ownership

* fix(release): repair web sync recovery gates

* fix(release): bind sync artifacts to validated base

* fix(release): verify synchronized file mode

* fix(release): paginate bot PR discovery
2026-08-11 18:37:49 +08:00
54b9cd8389 feat(opengateway): free retirement — paid Ling id, dual Nemotron, Macaron Venti (#2108)
The gateway retired its free models on 2026-08-10, keeping Nemotron 3
Ultra :free as the one free model (rate-limited by OpenRouter's shared
pool) and adding its paid throttle-free sibling as a separate entry.
Ling 3.0 Flash moves to its paid id (the :free id is aliased
server-side for older clients), Macaron V1 Tall is now paid, and
Macaron V1 Venti (748B MoL on GLM-5.2, 1M ctx) joins the catalog. The
ling entry id keeps its historical -free suffix so saved selections
resolve; HY3's stale Free label removed.

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-08-10 22:16:49 +08:00
JATMNandGitHub 93dbc72cbd refactor(openai-shim): extract Codex dispatch (#2074)
* refactor(openai-shim): extract Codex dispatch

* test(openai-shim): cover Codex dispatch guards
2026-08-10 11:10:02 +08:00
JATMNandGitHub 7fae0ffee0 refactor(openai-shim): extract request preparation (#2073)
* refactor(openai-shim): extract request preparation

* test(openai-shim): move request preparation regressions

* fix(openai-shim): rebase request preparation extraction

* test(openai-shim): assert converted preparation tools
2026-08-10 10:10:36 +08:00
95409464f3 feat(codex): move codexplan default to GPT-5.6 Sol (#2051)
* feat(codex): default codexplan to GPT-5.6 Sol

Preserve existing reasoning and routing behavior while updating the default model, labels, documentation, and focused coverage.

* test(codex): lock fallback and routing behavior

* make credential-step test self-contained

* fix(codex): default unset teammate fallback to GPT-5.6 Sol

The changes to the inert codex config keys are known to have no functional effect, but we updated them to ensure the defaults are correct and consistent across the table.

* fix codexplan gateway defaults after model resolution

* Revert "fix codexplan gateway defaults after model resolution"

This reverts commit 2b5ae791f8.

* fix codexplan custom gateway reasoning default

* fix codexplan reasoning query request routing

* fix(test): unmount CodexCredentialStep render instance before cleanup

---------

Co-authored-by: jatmn <the@jat.mn>
2026-08-10 10:09:48 +08:00
0xfandomandGitHub eb1de5b576 fix(bash): convert BRE interval braces when previewing sed edits (#1955)
* fix(bash): convert BRE interval braces when previewing sed edits

convertBrePatternToJs rewrites a POSIX BRE sed pattern into a JS regex so the
permission dialog can preview what a 'sed -i s/.../.../' edit will do. It
unescapes the BRE metacharacters that flip escaping between BRE and JS, but the
membership sets listed only +?|() and omitted { }. In BRE '\{n,m\}' is the
interval quantifier and a bare '{' is literal — the reverse of JS — so both
were handled backwards: '\{2\}' was emitted as a JS literal (matched nothing)
and a bare '{2}' became a JS quantifier.

The result is a misleading diff at an approval gate: 'sed -i s/a\{2\}/X/'
previews as no change while the command the user approves rewrites the file.
Add { } to both sets. Add coverage for escaped intervals, escaped ranges, and
bare-brace literals.

* fix(bash): only unescape sed interval braces when they form a valid count

\{,m\}, \{\} and \{n,m,k\} are literal brace runs in sed, not interval
quantifiers. Emitting them as JS braces produced a bogus quantifier (or a
literal that matched nothing), so the sed-edit preview showed no change while
real sed rewrote the file. Detect the enclosed count and only convert legal
BRE intervals (n, n,, n,m, and the GNU ,m extension, normalized to {0,m});
otherwise keep the braces escaped as literals. Extract the shared metachar set
and cover ? | ( ) and the ERE brace path.

* fix(bash): decline to simulate sed edits that cannot be reproduced faithfully

The preview is what the user approves, so it must either match what sed writes
or not be rendered at all. Three cases could not match:

- Zero-minimum quantifiers under g. sed and JS advance differently past an empty
  match: s/a\{0,3\}/X/g turns aaaab into XXbX in sed but XXXbX in JS, and the
  same holds for the pre-existing s/a*/X/g (XbX vs XXbX).
- Bracket expressions. \{ and \} are ordinary members inside [...], not an
  interval, and a backslash is a literal member in POSIX brackets but an escape
  in a JS character class, so [\{,3\}] cannot be mapped across.
- Illegal interval bodies. sed rejects \{\} and \{1,2,3\} outright (Invalid
  content of \{\}) and leaves the file untouched, so they are not literal braces
  and there is no edit to show.

parseSedEditCommand now returns null for these, falling back to ordinary bash
rendering. The same gate catches patterns whose translation is not a valid JS
regex, which previously threw and was swallowed into a silent no-change diff.

Verified differentially against GNU sed 4.10: of 22 expressions, the 15 still
simulated match sed byte-for-byte and the 7 divergent ones now decline.

* fix(bash): restrict the sed preview to portable, per-line-faithful patterns

Tighten the faithful-simulation gate on every axis where the preview could
disagree with what sed writes on the user's platform:

- Apply the substitution once per line, as sed does: without g, sed replaces
  the first match on EVERY line, so a whole-buffer replace previewed
  s/a\{2\}/X/ on 'aa\naa\n' as 'X\naa\n' where sed writes 'X\nX\n'. This also
  makes ^ and $ anchor per line, matching sed.
- Decline GNU-only operators: \+ \? \| and the \{,m\} interval are extensions
  that BSD/macOS sed treats as literals (or rejects outright), and this parser
  explicitly supports macOS through its -i '' handling — one platform's
  operator is the other's literal, so no single preview can be right for both.
  Alternation is additionally unfaithful even on GNU: POSIX selects the
  leftmost-longest branch where JavaScript takes the first that matches.
- Decline unterminated bracket expressions (an error in sed, not a literal)
  and POSIX [:class:]/[=equiv=]/[.collate.] constructs, which JavaScript would
  silently read as plain character sets. ERE patterns, which previously carried
  over verbatim, are screened for the same constructs.

Verified differentially against GNU sed 4.10 and macOS BSD sed across 27
expressions including multi-line inputs: every still-simulated pattern matches
both implementations byte-for-byte, and every declined pattern demonstrably
differs between platforms, differs from JavaScript, or errors in sed.

* fix(bash): keep empty files empty in the per-line sed simulation

An empty file has no lines, so sed never runs the substitution and the
output stays empty. Splitting '' fabricated one empty line, letting
anchored patterns like s/^/X/ preview an edit the real command does not
make. Return early before splitting; verified against GNU sed 4.10 and
BSD/macOS sed (both leave the file at 0 bytes).

* fix(bash): only simulate the sed subset that translates faithfully

The preview replaces the command once approved, so every accepted pattern
must produce exactly what sed writes. Move from screening known-bad
constructs to admitting only ones with matching semantics:

- flags: accept g/i/I only. 1-9 select the Nth match per line (the
  simulator always rewrites the first, so s/a\{2\}/X/2 on "aaaa" previewed
  "Xaa" where sed writes "aaX"); p prints; m/M redefine ^ and $.
- replacements: require literal text. \1-\9, &, \n and \U are sed syntax the
  simulator passes through verbatim, so s/\(a\)\{2\}/\1/ wrote the literal
  characters \1 where sed writes a.
- escapes: allow only the portable set. \< and \> are word boundaries in GNU
  sed but literal angle brackets in JS; \d is the converse.
- anchors: bare ^ and $ only anchor at a BRE boundary, so s/a^b/X/ previewed
  no change while sed rewrote the literal text.
- ERE intervals: validate the body there too — JS reads a{,3} as literal
  braces while GNU sed -E rewrites "aaaab" to "XXbX".
- character semantics: compile with u, so a quantifier counts characters as
  sed does rather than UTF-16 code units.
- CRLF: compile with s, so . matches the carriage return the pattern space
  holds.
- an empty pattern declines: sed has no previous regexp to reuse and errors.

Verified against GNU sed 4.10 and BSD/macOS sed: all 12 accepted expressions
match both byte-for-byte.

* fix(bash): decline bracket expressions opening with a ] member

POSIX treats the first ] in [] ] / [^] ] as an ordinary member, so GNU sed
rewrites "a]b" to "aXb". JavaScript reads it as the class terminator and
matches nothing, so the preview showed no change while sed rewrote the file.
findBracketEnd already skipped the leading ] to locate the real terminator,
but the body was then carried into the JS regex verbatim.

Decline on both the BRE and ERE paths.

* fix(bash): close three sed preview divergences before approval

All three let an approved preview differ from what the command writes,
which is the failure this simulator exists to avoid -- the preview is
persisted directly once the user approves it.

Dollar tokens: $ is an ordinary character in a sed replacement but a
substitution token to String.replace, so s/\(a\)/$1/ previewed the matched
text where sed writes the two characters $1. Double each $ for the JS
replacement.

ERE escapes: the ERE screen skipped every backslash escape and then used
the source verbatim, so -E 's/\d/X/g' on "1d2" previewed "XdX" while sed
writes "1X2" -- GNU sed reads \d as a literal d. \w, \s and \u{...} diverge
the same way. Admit only the escapes that are the same literal in both
dialects, mirroring the BRE allowlist.

Case-insensitive matching: the emitted regex needs u for the quantifier
fix, and u + i selects ECMAScript Unicode case folding rather than sed's
locale matching, so s/k/X/I rewrote a Kelvin sign that GNU sed under
C.UTF-8 leaves alone. Decline i/I until that can be modeled.

Verified the accepted set against GNU sed 4.10: 11 expressions, including
each dollar case above, byte-identical.

* fix(bash): gate the sed preview on locale and bound its matching

Three follow-ups to the preview-fidelity work.

The emitted regex always carries the u flag so a quantifier counts
characters, but sed inherits the process locale and counts bytes in a byte
locale: LC_ALL=C 's/.\{2\}/X/' on an emoji consumes two of its bytes and
leaves the rest in the file. Claim a sed edit only when the resolved locale
(LC_ALL, then LC_CTYPE, then LANG, per POSIX) names a UTF-8 codeset.

Interval support also made nested quantifiers translatable:
\(a\{1,\}\)\{1,\}b becomes (a{1,}){1,}b, which backtracks
exponentially on a run of a's with no b. applySedSubstitution runs
synchronously while the permission request renders, so that stalls the
approval UI before the user can decide. Decline a quantified group that
already contains a quantifier.

The replacement gate rejected every backslash, including \/ and \&, which
the translation right below already handles faithfully -- so ordinary
commands like s/foo/path\/to/ lost their file diff for no reason. Admit
those two, keep declining backreferences, case folding and a bare &.

Verified against GNU sed 4.10: the newly readmitted escapes and the
still-accepted single-quantifier groups all match byte-for-byte.

* fix(bash): decline a repeated s/// g flag the way sed does

GNU sed rejects a duplicated flag ("multiple `g' options to `s' command"),
so tighten the accepted-flags pattern from /^g*$/ to /^g?$/. A preview that
rendered `s/a/X/gg` as a successful global rewrite would diverge from the
command sed refuses to run.

* test(bash): scope the sed preview test to the CRLF-normalized gate

The permission path normalizes CRLF to LF before calling applySedSubstitution,
so the approved preview never sees a raw carriage return. Replace the
raw-\r\n assertion (which claimed a fidelity the gate does not exercise) with
one over LF content, and document that raw-CR bytes are out of scope.

* fix(bash): decline ERE (?...) groups sed does not implement

ereHasUnfaithfulConstructs screened escapes, alternation, anchors, intervals
and bracket bodies but treated grouping as implicitly safe. POSIX/GNU sed -E
only supports plain capturing (...); a (? opens JavaScript-only syntax --
(?:), lookaround, named groups -- that new RegExp compiles but GNU sed rejects.
The preview would render a concrete edit for a command sed refuses to run, so
decline as soon as an unescaped ( is followed by ?.
2026-08-10 10:09:00 +08:00
0xfandomandGitHub 41d2f3b831 fix(repomap): resolve file language by real extension, own-property only (#2100)
getLanguageForFile computed the extension as substring(lastIndexOf('.')). For a
path with no dot, lastIndexOf returns -1 and substring(-1) clamps to
substring(0), so the whole filename became the lookup key against the plain
SUPPORTED_EXTENSIONS object. That misclassifies extensionless files, and a root
file named after an Object.prototype member (constructor, __proto__, toString,
…) resolves to an inherited value that the ?? null guard accepts as a supported
language -- so isSupportedFile returns true and the file enters the repo-map
graph with a bogus language. Return null when there is no dot and gate the
lookup on Object.hasOwn.
2026-08-07 23:33:30 +08:00
JATMNandGitHub deb91941e1 refactor(openai-shim): extract response adapters (#2072)
* refactor(openai-shim): extract response adapters

* test(openai-shim): cover response adapter stream wrappers

Add focused regression tests for geminiSseToAnthropic and
openaiStreamToAnthropic through the responseAdapters facade wiring.

Validated with: bun test src/services/api/openaiShim/responseAdapters.test.ts

* test(openai-shim): assert Gemini tool-use stream blocks in adapter test

Extend the responseAdapters geminiSseToAnthropic wrapper test to cover
tool_use content_block_start, input_json_delta, and content_block_stop.
Remove stale post-extraction imports from the openaiShim facade.

* test(openai-shim): cover facade parser re-exports

Add a focused openaiShim.test.ts case that imports parseTextToolCalls and
parseXmlToolCalls through the public facade and asserts shared sequencing.
2026-08-07 23:32:31 +08:00
0xfandomandGitHub c327805e1d fix(cost): guard model-cost lookup against prototype-member model ids (#2064)
* fix(cost): guard model-cost lookup against prototype-member ids

MODEL_COSTS is a plain object, so `MODEL_COSTS[shortName]` inherits
Object.prototype members. A model id of `constructor` or `__proto__`
-- both valid arbitrary ids for custom/OpenAI-compatible providers, and
already lowercase so getCanonicalName returns them unchanged -- resolved
to a truthy prototype value (the Object constructor / Object.prototype),
so the `!costs` unknown-model guard was skipped: trackUnknownModelCost
never fired and tokensToUSDCost read undefined rate fields, producing a
NaN cost that permanently poisons the running session total (total + NaN
stays NaN) and surfaces as "$NaN". getModelPricingString had the same
defect and rendered "$NaN/$NaN per Mtok".

Match on own properties via Object.hasOwn, mirroring resolveOutputStyle
in constants/outputStyles.ts. Add a regression test asserting proto-name
ids take the unknown-model path and yield a finite, positive cost.

* fix(cost): guard per-model usage tracking against prototype-member ids

getModelCosts was hardened, but the sibling per-model accounting kept the
same latent hole. STATE.modelUsage is a plain object, so a model id of
`constructor` / `__proto__` (arbitrary for custom/OpenAI-compatible
providers) reaches both getUsageForModel's read and the
`STATE.modelUsage[model] = ...` write. On the read, an absent proto-name
key resolves to an inherited Object.prototype member; addToTotalModelUsage
then does `modelUsage.inputTokens += ...` on that inherited object, and for
`__proto__` that lands on Object.prototype itself -- process-wide pollution
(every object gains inputTokens = NaN, etc.). On the write, bracket-setting
`__proto__` invokes the prototype setter.

Back modelUsage with a null-prototype map (emptyModelUsage) at init, reset,
and restore, so both operations act on ordinary own keys, and read through
Object.hasOwn to match the getModelCosts guard. Restore re-keys a persisted
breakdown (which can carry an own `__proto__` from JSON) into the null-proto
map. Regression test asserts no Object.prototype pollution and correct own-key
round-trip for proto-name ids.

* test(cost): tighten the proto-name model-cost regression

Assert Number.isFinite(cost) (rejects Infinity too, not just NaN) and that
the unknown-model detection flag fires, so a regression that dropped
trackUnknownModelCost while keeping the fallback tier would fail. Reset the
process-wide cost state afterward to avoid leaking into other suites, and
correct the comment: getModelPricingString has no production callers and
pre-fix threw a TypeError rather than rendering "$NaN/$NaN per Mtok".

* fix(cost): guard the /cost per-model aggregation against prototype-member ids

formatModelUsage accumulates per-model usage into a plain object keyed by
canonical short name. An unrecognised custom-provider id that canonicalizes to
`__proto__` / `constructor` (unchanged, since it matches no Claude pattern) made
the `!usageByShortName[shortName]` check read an inherited Object.prototype
member, skip initialization, and increment it in place -- for `__proto__` that
mutation lands on Object.prototype process-wide, and the model is dropped from
the displayed /cost breakdown. Use a null-prototype accumulator and an
Object.hasOwn guard, matching the getModelCosts / getUsageForModel fixes.

Regression drives the real addToTotalSessionCost -> formatTotalCost path for
`__proto__` and `constructor` ids, asserting no prototype pollution and that
both appear in the breakdown.

* test(cost): tidy the /cost proto-pollution regression harness

Load addToTotalSessionCost via a lazy ESM import instead of require (drops the
eslint suppression) and snapshot the guarded Object.prototype descriptors so
cleanup restores pre-existing state instead of unconditionally deleting keys a
sibling module might legitimately own.
2026-08-07 09:58:21 +08:00
BogdanandGitHub d834904e5a fix(session): make transcript replacements crash-safe (#2094)
* fix(session): make transcript replacements crash-safe

Complete transcript rewrites could truncate live JSONL files before preserved data was durable, risking unrecoverable resume history after an interrupted write. Commit replacements through exclusive sibling temp files and serialize them with all transcript append paths so readers observe either the old file or the complete replacement.

* fix(session): preserve concurrent transcript updates

Abort tombstone commits when the scanned transcript changes before replacement, and keep existing local history when remote foreground hydration returns no entries. Harden the associated portability, option coverage, queue timing, and diagnostics.

* test(session): match hydration reader signature

Pass the explicit optional subagent reader in the empty-hydration regression so a fresh TypeScript build sees the complete helper signature.

* fix(session): coordinate transcript writers across processes

Hold a same-directory cooperative lock across transcript replacement and final-line truncation, and make session plus SDK append paths participate. Exercise the post-validation/pre-rename race deterministically so external appends land after the complete commit.

* test(session): provide empty hydration subagent reader

* fix(session): scope transcript lock ownership

Separate async and synchronous lock ownership so unrelated sync appends cannot bypass an in-flight replacement. Route aliased in-process appends through the queue, propagate lock compromise through AbortSignal, and cover both symlink-alias and rename-boundary races.
2026-08-07 09:57:32 +08:00
BogdanandGitHub 6465a516f2 fix(mcp): serialize OAuth and XAA refresh across processes (#2093)
* fix(mcp): serialize OAuth and XAA refresh across processes

Normal OAuth refresh, reactive 401 recovery, and silent XAA exchange can otherwise race shared secure-storage writes between processes. Coordinate them on one server-scoped lock and re-read storage so waiters reuse persisted winners.

* fix(mcp): harden refresh follow-up paths

Use asynchronous cache-bypass reads on request paths while preserving the adjacent final record merge and write. Make the XAA concurrency fixtures independent of module import order and extend abort, redaction, and retry coverage.

* fix(mcp): honor aborts after credential reads

Check the active cancellation signal after asynchronous secure-storage reads so fresh-token fast paths cannot return credentials to an aborted request. Cover cancellation while a cache-bypassing read is pending.
2026-08-07 09:56:12 +08:00
BogdanandGitHub d427a4b2bb perf(cli): enable Node module compile cache (#2092)
* perf(cli): enable Node module compile cache

Warm CLI invocations spend substantial time compiling the bundled ESM entrypoint. Enable Node's optional on-disk compile cache only in the process that imports the bundle, while preserving early Node 22 compatibility and making cache failures non-fatal.

Add deterministic launcher coverage, packaging checks, and a reproducible benchmark procedure so the startup benefit can be measured without flaky CI thresholds.

* fix(ci): isolate minimum Node launcher check

The full validation suite depends on knip and oxc-parser behavior unavailable in Node 22.0.0. Keep full CI on the active Node 22 line and exercise the declared runtime floor in a dedicated build-and-launch job.

* fix(benchmark): harden startup measurements

Keep environment setup outside the timed process window, document the API's Node 22.8 floor, and preserve completed benchmark results when git metadata is unavailable.

* test(cli): verify compile cache disable behavior

Pair NODE_DISABLE_COMPILE_CACHE with a temporary cache directory and assert that supported Node releases leave it empty while preserving normal launcher output.
2026-08-07 09:55:01 +08:00
JATMNandGitHub bac012f70b refactor(openai-shim): extract transport lifecycle (#2071)
* refactor(openai-shim): extract transport lifecycle

* refactor(openai-shim): rebase transport extraction and address review

Rebase onto current main and clean up the transport extraction follow-ups:
drop stale facade imports left after the move and restore the full API
timeout parser negative-case coverage in transport.test.ts.

* test(openai-shim): address CodeRabbit review findings

Use path.join in the architecture guard, require transport.ts in the
mandatory extraction slice, strengthen Gemini stream conversion coverage,
and add transport deadline/cancellation regression tests with fake timers.

* test(openai-shim): assert manual signal cleanup after body cancel

Exercise the combineRequestSignals fallback without AbortSignal.any so
early body cancellation removes caller listeners and a later caller.abort
does not abort the combined fetch signal.

* test(openai-shim): restore AbortSignal.any when initially absent

Delete the temporary AbortSignal.any override when the runtime did not
define an own property, so transport and facade signal-cleanup tests leave
global AbortSignal state unchanged for later cases.
2026-08-07 09:51:48 +08:00
BogdanandGitHub 1bf8076d48 fix(input): preserve text in DEL-coalesced chunks (#2091)
* fix(input): preserve text in DEL-coalesced chunks

Some terminal transports deliver replacement input as raw DEL bytes and printable text in one read. The raw-DEL workaround previously applied only the deletions and returned, dropping the replacement text and leaving same-event cursor and mode state stale.

Process filtered chunks in source order through the existing cursor semantics, preserve coalesced submission and Vim state, and cover grapheme, token, filter, mode, and batching cases.

* test(input): harden DEL regression coverage

* test(input): clean up harnesses after timeouts

* fix(input): preserve coalesced consumer state

* fix(input): synchronize coalesced mode state
2026-08-07 09:49:01 +08:00
மனோஜ்குமார் பழனிச்சாமிandGitHub 95eeb0bde3 feat(cli): add --yolo alias for --dangerously-skip-permissions (#2097)
Register the alias on the main command and the ssh stub. Recognize it in
the cc:// and ssh raw-argv scans, and in both skills pre-parse boolean sets
(leading and trailing), so  and
 route correctly. Update the web flags docs.

Includes source-scan + help-text tests proving the alias is wired through.
The SSH/argv refactor remains on the existing feat/yolo-flag branch for a
separate follow-up PR.
2026-08-07 09:47:47 +08:00
JATMNandGitHub b0cbfe1100 fix(repl): make local interactive max-turns configurable (#2086)
* fix(repl): make interactive max-turns configurable

Wire --max-turns into interactive sessionConfig and honor
OPENCLAUDE_MAX_TURNS / CLAUDE_CODE_MAX_TURNS so long autonomous REPL
sessions can raise the default 50-turn per-prompt cap (fixes #2079).

* fix(repl): forward --max-turns on connect/ssh/remote launches

sessionConfig covered the normal interactive paths; connect, SSH,
assistant, and --remote built REPL props without spreading it, so the
CLI override was dropped despite help advertising interactive support.

* fix(repl): scope interactive max-turns to local query loops

Remote-backed sessions bypass local query(), so forwarding --max-turns
into those REPL props over-claimed enforcement. Clarify help/docs and
match OPENCLAUDE_MAX_RETRIES precedence when OPENCLAUDE_MAX_TURNS is set
but invalid.

* feat(config): add interactive max turns under /config

Expose replMaxTurns in the Config panel (50/100/200/500) and resolve it
after CLI/env so local interactive sessions can raise the per-prompt
cap without restarting. Resolve at query time so mid-session /config
changes apply on the next prompt.

* docs(repl): clarify invalid OPENCLAUDE_MAX_TURNS precedence

Match the OPENCLAUDE_MAX_RETRIES contract: a set-but-invalid primary
env var uses the default and does not fall through to legacy or /config.

* fix(repl): address PR review on max-turns help and web version gate

Share the --max-turns Commander description via an imported constant so
help and tests stay in sync without breaking the CLI bundle, replace
source-only help assertions with Commander behavior coverage, and add
the published 0.27.0 entry so web verify-dist passes.

* fix(repl): typecheck Commander maxTurns opts and warn on invalid env

Avoid TS2339 on untyped Commander opts, log invalid OPENCLAUDE_MAX_TURNS
like MAX_RETRIES, and clarify that /config shows the persisted preference.

* fix(repl): warn when max turns is unlimited

* fix(repl): scope unlimited-turn warning locally

* fix(repl): preserve interactive turn caps across backgrounding

* fix(repl): preserve turn caps when backgrounding

* fix(repl): share turn budget across background handoff

* fix(repl): reserve turns at provider dispatch

* fix(repl): snapshot background handoff transcript

* fix(tasks): avoid phantom background session task

* fix(repl): preserve handoff lifecycle state

* fix(repl): own pending background handoffs

* test(tasks): isolate background session task storage

* fix(repl): refresh background task title and test cleanup

* test(repl): cover max-turn CLI dispatch paths

* test(queue): cover prepend notification and priority

* fix(repl): skip background handoff after foreground query throws

Rebased onto main and gate Ctrl+B continuation on !didThrow so a faulted
foreground turn cannot start a background session from partial state.

* fix(repl): address PR review findings on notifications and handoff

Dedupe background task notifications by embedded task id, gate background
continuation on preflight veto, scope queue removal to main-thread notifications,
and forward maxTurns through all launchRepl entry points.

* fix(repl): resolve latest CodeRabbit inline review findings

Dedupe claimed notification batches by task id, restore notifications on
pre-registration abort, tighten test isolation, and replace remaining brittle
source-text assertions with behavioral coverage.

* fix(repl): keep notification restore active until provider dispatch

Stop clearing notification ownership when preparation succeeds so pre-dispatch
aborts can restore claimed queue items, and commit ownership once the provider
starts. Add regression coverage for the abort path and headless max-turns zero.

* test(repl): cover post-dispatch ownership and headless max-turns 0

Add regression tests for notification restore after provider dispatch commits
ownership, and assert headless --max-turns 0 reaches query() without interactive
resolution stripping the value.

* fix(repl): restore only embeddable notifications on Ctrl+B handoff abort

Track the deduped successor subset when restoring claimed main-thread task
notifications so items already in the settled foreground transcript are not
re-queued. Clarify that agent-scoped notifications intentionally stay on their
owner drain path (issue #2079 scope is interactive turn caps only).

* fix(repl): address review findings on background handoff

Commit notification ownership when background sessions complete without
provider dispatch, forward all task notifications on Ctrl+B again, and
restore deferred max-turn cap attachments when continuation is cancelled.

* fix(repl): guard deferred cap restore and skip remote turn limits

Anchor deferred max-turn restoration to the handed-off transcript tail so a
cancelled Ctrl+B handoff cannot attach the prior prompt's cap to a newer turn.
Apply the interactive turn cap only in local sessions and align remote-session
docs/help wording.

* fix(repl): use messagesRef for deferred cap transcript anchor

persistentMessages is block-scoped inside onQuery try; read the settled
tail from messagesRef in finally so typecheck passes.

* test: harden context fallback warning assertion after max-turns tests

Scope the unknown-model context test to [context] warnings only so unrelated
import-time debug logs do not fail CI, and clear turn env vars in both that
test and replMaxTurnsProp setup to avoid cross-file pollution.

* test: address PR review findings on headless max-turns boundary

Add a runHeadless-to-ask regression that asserts maxTurns 0 is forwarded
through the headless print path, and restore OPENCLAUDE_MAX_TURNS env vars
in context.test.ts after the unknown-model fallback test mutates them.

* test: tidy headless max-turns boundary test and env isolation

Mock headless stdout so runHeadless completes cleanly without leaking
output, restore spies in finally, and centralize turn-env cleanup in
context.test beforeEach.

* fix: address PR review findings for max-turns background handoff

Separate model-request lifecycle from provider dispatch acceptance so
interruption correction arms before async prep, notification ownership
commits only after dispatch, deferred turn caps restore on every abort
path, and foreground work stays blocked while handoff preparation runs.
2026-08-06 20:24:34 +08:00
0xfandomandGitHub 2c42a325d9 fix(permissions): anchor the session plan-file match on its exact shape (#1994)
* fix(permissions): anchor the session plan-file match on its exact shape

isSessionPlanFile auto-allows the current session's plan file for both
read (checkReadableInternalPath) and un-prompted write
(checkEditableInternalPath). It matched with a bare
normalizedPath.startsWith(join(plansDir, planSlug)), which also accepts
any sibling whose name merely begins with the slug — {slug}nova.md,
{slug}-other.md, or a newly-created {slug}dir/ subtree. Those are not this
session's plan yet were silently readable and writable without a prompt.

Anchor on the two shapes getPlanFilePath actually emits: {slug}.md exactly,
or a {slug}-agent- prefix for subagent plans. Extract the decision into a
pure isPlanFilePath(plansDir, slug, path) helper so it can be unit-tested
without session state. normalize() still runs first, so traversal segments
can't escape the plans directory.

Same missing-separator class as the path-containment fix in #1974.

* fix(permissions): restrict the agent-plan branch to a single filename

The -agent- prefix check still matched any path beneath a lookalike
sibling directory: {plansDir}/{slug}-agent-evil/anything.md passed
startsWith and ended in .md, so both permission carve-outs granted
unprompted read and write to arbitrary files below it. The malformed
{slug}-agent-.md, which getPlanFilePath never emits, was accepted too.

Require the remainder after the prefix to be exactly one nonempty agent id
followed by .md — no path separators.

* fix(plans): keep separator-carrying agent ids in one filename component

The anchored predicate rejected any agent id containing a path separator,
but producers can emit one: TeamCreateTool accepts any nonblank team name
and teammate spawning only strips `@` from the teammate name, so a team
called `a/b` yields the path {plansDir}/{slug}-agent-writer@a/b.md. That
is a file in a subdirectory, not a plan file, so the teammate lost the
carve-out for its own plan and was blocked in plan mode.

Escape the separators where the path is built instead. Percent-escaping is
reversible, so two teammates can never collide on one plan file, and ids
without those characters are untouched -- existing plan files keep their
paths.

* fix(plans): recover plans written under the unescaped agent id

Escaping changes the pathname for teammates whose id already contains a
separator, and team names have always accepted arbitrary nonblank text --
so plans for ids like writer@a/b or writer@100% are already on disk under
the raw name. Every reader now builds the escaped name, so on upgrade the
teammate's plan reads as missing and a second file is created beside it.

getPlan falls back to the unescaped path on ENOENT and moves the file to
the escaped name. Moving rather than copying is what makes it stick: the
escaped name is the one the permission carve-out recognizes, so a plan left
at the old path would keep falling through to ordinary permission handling
on every later write. A failed move is not fatal, the content is already
read.

The recovery takes explicit paths so it is covered against a real temporary
directory rather than a mocked filesystem.

* fix(plans): confine legacy plan recovery to the plans directory

readLegacyUnescapedPlan builds the pre-escape path from the raw, unescaped
agent id so an existing file can be found. Team/agent names accept arbitrary
nonblank text, so a traversal-shaped id (`../../../etc/passwd`) collapses to
a path outside the plans directory -- which readAndMigrateLegacyPlan then
reads and renames, moving an arbitrary file. Refuse any resolved path the
plans directory does not contain before delegating.

* fix(plans): give escaped agent plans a collision-free namespace and harden recovery

The escaped filename shared a directory with legacy plans, so two distinct
teammates could map onto one file: `writer@a/b` writes the escaped
`{slug}-agent-writer@a%2Fb.md` while `writer@a%2Fb` already owns that exact
name as its raw legacy plan -- a cross-agent read and clobber. Store escaped
agent plans under a dedicated `agents/` subdirectory: a real path separator
is the one thing a raw single-component legacy name can never contain, so
the two namespaces are provably disjoint. The permission carve-out
(isPlanFilePath) recognizes the new location.

Harden legacy recovery, which reads then renames a file built from the raw
(unescaped) agent id:
- Reject any `..` segment before building the path, so `a/../{slug}` can no
  longer collapse onto the main plan (or `a/../{slug}-agent-victim` onto a
  sibling) and have recovery move another agent's file.
- Make migration no-clobber: never rename a legacy file over a plan already
  present at the escaped path.
- Export readLegacyUnescapedPlan (with injectable plansDir/slug) so the guard
  is covered through the recovery flow, not just isPathWithinPlansDir alone.

* test(plans): cover getPlan's ENOENT recovery wiring end to end

The recovery helpers are unit-tested, but nothing drove getPlan() itself
through the ENOENT fallback -- the whole user-visible fix. Add a test that
plants a legacy plan under a temp config dir and asserts getPlan() returns
its contents and migrates it into the agents/ subdirectory, serialized
under the shared mutation lock since it swaps OPENCLAUDE_CONFIG_DIR.

* fix(plans): anchor plan-file matching on the canonical encoding and harden recovery

Addresses review on the agent-plan permission carve-out.

isPlanFilePath accepted any `{slug}-agent-<x>.md` whose `<x>` had no raw
`/` or `\`, but getPlanFilePath emits only the canonical output of
encodeAgentIdForPlanFile (escapes `%`->`%25`, `/`->`%2F`, `\`->`%5C`). So a
raw-percent sibling such as `{slug}-agent-writer@100%.md` (canonical form
`...writer@100%25.md`) was auto-allowed for unprompted read/write even though
the producer never writes it. Add decodeAgentIdForPlanFile and
isCanonicalPlanFileEncoding (a component is canonical iff re-encoding its
decode reproduces it byte-for-byte) and anchor the agent branch on it. This
accepts every path the encoder can emit and rejects raw-`%`/raw-separator
lookalikes, subsuming the previous separator-only check.

Also harden legacy recovery, which reads and renames a path built from the
raw agent id:
- getPlan now treats an empty/whitespace escaped file as not-a-plan and falls
  through to legacy recovery. isPlanFilePath permits a direct FileWrite/FileEdit
  to the canonical escaped path before migration runs; such a stub would
  otherwise permanently shadow a legacy plan that still holds content. Recovery's
  no-clobber guard returns the legacy contents without renaming over the stub,
  so a genuine concurrent escaped write is never lost.
- readAndMigrateLegacyPlan lstat-checks the legacy slot and refuses anything
  that is not a regular file, so a symlink planted there cannot make recovery
  read and rename an arbitrary target outside the plans directory.

Tests: canonical-vs-lookalike pairs for `%`/separator ids, getPlan driven
end-to-end for a separator id and for the empty-stub fallthrough, and a
symlinked legacy slot. The getPlan integration tests acquire the shared
mutation lock inside try/finally and clear the plan slug on teardown.

* fix(plans): close symlink and race gaps in plan-file recovery and the carve-out

Second review pass on the agent-plan hardening.

- Symlinked path components no longer bypass the lexical carve-out. The plan-file
  permission grant (isSessionPlanFile) now resolves the deepest existing ancestor
  of the target and requires it to stay within the *resolved* plans directory, so
  a symlinked `agents` subdir (or plans dir) that redirects the real file outside
  the plans directory is refused instead of auto-allowed. Legacy recovery gets the
  same containment check, closing the slash-bearing-id case where a symlinked
  intermediate `{slug}-agent-writer@a` parent passed the prefix checks and leaf
  lstat.
- Migration is now a genuine no-clobber move: linkSync (atomic, fails EEXIST)
  replaces the existsSync-then-renameSync check-then-act race that could replace a
  concurrently-created live plan on POSIX. The escaped hard link pins the inode we
  lstat'd, and we read through it, so a symlink swap of the legacy pathname cannot
  redirect the read. Reads verify the inode/device are unchanged across the read.
- Traversal validation uses the host platform's real separators: on POSIX `\` is a
  legal filename character, so a legacy id like `a\..\b` (persisted as one flat
  filename) recovers again instead of being wrongly rejected; Windows still treats
  both `/` and `\` as separators.

Tests: symlinked intermediate directory rejection (helper + recovery), POSIX
literal-backslash recovery, genuine-move semantics. All fail on the pre-fix code.

* refactor(permissions): reuse the shared plans-dir containment helper

Drop the duplicate isResolvedWithinPlansDir in the permission layer and route
the session plan-file carve-out through the exported isResolvedPathWithinPlansDir
from plans.ts, keeping the symlink-containment logic in one place. Guard the two
symlink-based tests on non-Windows so they skip where symlinkSync needs
privileges.
2026-08-06 20:23:21 +08:00
keyarvsfandGitHub 248424ffe3 fix(model-picker): eliminate O(n²) catalog rebuild lag in /model (#2078)
* fix(model-picker): eliminate O(n²) catalog rebuild lag in /model

getModelOptions() ran an O(n²) optionMatchesModel loop (catalog scan per
option) plus an O(n²) duplicate-apiName filter, costing ~43ms per call on
catalogs with hundreds of models (e.g. Fireworks' ~280 entries). The
picker also rebuilt the full options list on every keystroke via
isGenuineSwitchProfileValue, so arrow-key navigation lagged badly.

- hoist catalog lookup out of the per-option loop (hasOptionValue)
- precompute duplicate apiNames into a Set
- short-circuit isGenuineSwitchProfileValue for non-switch values

getModelOptions(): 43.5ms -> 2.0ms on a 277-entry catalog

* perf(model-options): share route-catalog context across getModelOptions checks

getRouteCatalogModelOption re-resolved the active route, fetched the
catalog entries, and rebuilt the duplicate-apiName set on every call —
getModelOptions() invoked it up to 3x per build (env custom model, each
scoped additional option, active custom model + fallback), so large
catalogs (Fireworks ~280 entries) paid O(n) context rebuilds repeatedly.

Build the RouteCatalogContext lazily once per options build and pass it
through findRouteCatalogOption/hasOptionValue. Behavior unchanged; the
catalog-miss path measures 6.1ms -> 4.1ms on a 277-entry catalog.

* test(model-picker): add regression tests for catalog dedup and switch-profile guard

* test(model-picker): gate process-wide mocks in regression tests

* test(model-picker): reuse mocked modelOptions instance in switch-profile test

* test(model-picker): prevent mock leakage

* test(model-picker): drop dead env overrides overwritten by catalog dedup helper

OPENAI_BASE_URL and OPENAI_MODEL assigned in the scoped-cache test were
immediately overwritten by getRouteCatalogModelOptions, so they never
affected the exercised path. Remove the dead assignments; keep
OPENAI_API_KEY to preserve the helper's auth path.

* test(model-picker): assert getModelOptions skipped for ordinary ids

isGenuineSwitchProfileValue short-circuits on the switch-profile prefix,
skipping the getModelOptions() rebuild for ordinary model ids. Track the
gated getModelOptions binding ModelPicker captures (opt-in call-through
mock in importFreshModelPicker) and assert it is never invoked for
non-prefixed ids, alongside the existing false-result assertions.

* test(model-picker): address review feedback on spies, fetch bounds, and alias dedup
2026-08-06 20:22:29 +08:00
5844d1fe8a Feat/ultracode blue spinner (#2096)
* feat(ultracode): add blue/cyan spinner and effort visual treatment

- Add EFFORT_ULTRACODE (◆) figure for effort display surfaces
- Add ultracode/ultracodeShimmer theme colors across all 6 theme variants
- Wire ultracode case into effortLevelToSymbol() for icon rendering
- Use blue-cyan RGB shimmer for "thinking" text when ultracode is active
- Set spinner color override in REPL when displayed effort is ultracode

* feat(ultracode): tint prompt border and unify shimmer to theme tokens

Add a persistent cyan-blue prompt border whenever ultracode is the active
effort, reacting immediately to /effort and ranking below bash/teammate
overrides. Derive the spinner thinking-shimmer from the ultracode/
ultracodeShimmer theme tokens (with an ANSI/daltonized fallback) instead of
a divergent hardcoded cyan, so border, spinner, and shimmer share one source
of truth.

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

* test(spinner): cover ultracode shimmer color selection and ANSI fallback

Extracts the thinking-shimmer color computation into an exported
getThinkingShimmerColor helper (renderToString strips ANSI color, so the
selection logic is only observable through a direct call) and adds focused
tests for ultracode rgb() token interpolation, the ansi:* fallback
endpoints, and the non-ultracode gray interpolation.

Addresses CodeRabbit review on #2096.

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

---------

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-08-06 13:07:46 +08:00
JATMNandGitHub 63fda83d55 fix(web): add v0.27.0 changelog entry and clarify release-data ownership (#2075)
* fix(web): add 0.27 release entry

* test(web): cover 0.27 release entry

* fix(web): clarify 0.27 permission-timeout highlight

* docs: ban drive-by edits to web/src/data/releases.ts

Tell agents and contributors that the curated changelog is owned by
the release/web process, and stop verify-dist from instructing unrelated
PRs to patch it when npm publishes ahead of the site.

* test(web): locate curated release by version

* test(web): enforce newest-first release order
2026-08-03 10:54:52 +08:00
JATMNandGitHub b3735bedb3 refactor(openai-shim): extract request executor helpers (#2011)
* refactor(openai-shim): extract request execution

* fix(openai-shim): rebase executor extraction

* test(openai-shim): preserve local stream options coverage

* test(openai-shim): isolate Azure compatibility state

* fix(openai-shim): preserve executor retry contracts

* fix(openai-shim): retain route credential isolation

* fix(openai-shim): preserve executor transport behavior

* fix(openai-shim): avoid duplicate local retries

* fix(openai-shim): preserve LongCat credential routing

* fix(openai-shim): retain executor abort contracts

* fix(openai-shim): preserve fallback cancellation

* fix(openai-shim): preserve executor retry and transport contracts

* fix(openai-shim): stabilize extracted executor smoke coverage

* test(openai-shim): cover extracted executor retry contracts

* test(openai-shim): assert redacted HTTP errors

* test(openai-shim): isolate Azure executor configuration

* fix(openai-shim): preserve executor recovery retries

* test(openai-shim): remove migrated executor duplicates

* docs(openai-shim): clarify extracted façade budget

* test(openai-shim): enforce extraction modules

* fix(openai-shim): stop retrying cooled GitHub keys

* fix(openai-shim): avoid concurrent cooled-key retries

* test(openai-shim): stabilize pooled-key retry coverage

* fix(openai-shim): preserve newer credential cooldowns

* docs(openai-shim): explain stale auth eviction
2026-07-31 14:39:26 +08:00
github-actions[bot]GitHubgithub-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
7eeb90fb5b chore(main): release 0.27.0 (#2055)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
v0.27.0
2026-07-31 08:48:38 +08:00
5cac15cbda fix(minimax): mark MiniMax-M2.7 as text-only input (#2068)
Co-authored-by: octo-patch <266937838+octo-patch@users.noreply.github.com>
2026-07-30 21:52:02 +08:00
77c82829c4 docs(readme): add npm monthly downloads badge (#2069)
Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-07-30 15:57:28 +08:00
JATMNandGitHub 8df37c78f4 fix(agents): allow subagents from multi-repo parent sessions (#2063)
* fix(agents): allow subagents from multi-repo parent sessions

Expose Agent cwd in the open build, let cwd select the child repo for
worktree isolation, and fall back instead of hard-failing when the
session itself is outside a git repository.

* fix(agents): persist cwd on resume and forward it to worktree hooks

Address final-head review: store explicit Agent cwd in metadata for
resume, pass cwd into WorktreeCreate hooks, reject relative cwd in the
schema, and make the multi-repo parent regression sandbox portable.

* fix(agents): keep child-repo cwd across worktree cleanup and resume

Persist explicit Agent cwd even when a worktree is created, preserve it
when unchanged worktrees are removed, and base fork worktree notices on
the child-repo cwd for multi-repo parent sessions.

* fix(agents): re-persist child-repo cwd on every resume

Always forward persisted Agent cwd through resume metadata writes so a
mid-life resume cannot drop the multi-repo fallback path, and tighten
prompt wording to match the missing-git fallback contract.

* fix(agents): validate cwd directories and recover from worktree hook failures

Require Agent cwd to be an existing directory, always re-persist the
original resume metadata cwd, and fall through from failed WorktreeCreate
hooks to git when the selected cwd is a git repository.

* fix(agents): keep WorktreeCreate hooks authoritative

Revert silent git fallback after hook failure. Treat WorktreeCreate hook
errors as recoverable in AgentTool so multi-repo cwd overrides still work
without bypassing configured hooks at the worktree layer.

* fix(agents): only soft-fallback missing-git worktree errors

Keep WorktreeCreate hook failures hard-failing so configured hooks stay
authoritative in normal git sessions. Soft-fallback remains limited to
the missing-git path that #2052 needs.

* docs(agents): clarify missing-git cwd fallback wording

Align AgentTool prompt and resume debug logs with the missing-git-only
soft-fallback contract for multi-repo parent sessions.

* fix(agents): keep fork worktree notices on session cwd

Inherited fork context paths are relative to the parent session, so the
worktree notice must use getCwd() even when isolation used a child-repo cwd.

* docs(agents): align runAgent cwd JSDoc with resume persistence

* fix(agents): address CodeRabbit cwd validation review notes

Use afterAll for schema-test temp cleanup, and preserve the underlying
stat failure reason when Agent cwd validation rejects a path.

* fix(agents): surface worktree isolation fallback visibly

Make the missing-cwd schema test path platform-neutral, and record a
user/model-visible notice plus tool-result flag when worktree isolation
soft-falls back outside a git repository.

* fix(agents): surface worktree fallback when sync agents background

Share async_launched payload construction so the sync-to-background path
includes worktreeIsolationFallback when worktree isolation soft-falls back.
2026-07-30 11:56:59 +08:00
JATMNandGitHub 871bf28568 refactor(openai-shim): extract typed request body planning (#2010)
* refactor(openai-shim): extract request planning

* test(openai-shim): cover Bankr route credential base

* test(openai-shim): retain empty Responses fallback coverage

* fix(openai-shim): address planner review findings

* test(openai-shim): cover serializeBody transport routing

Add planner tests that exercise serializeBody() for responses,
anthropic_messages, and gemini transports, including omit-flag rebuilds.
Remove the vacuous Gemini tool_choice assertion that never guarded behavior.
2026-07-30 11:55:24 +08:00
e636f7d1cb feat(opengateway): add Macaron V1 Tall to the gateway catalog (#2067)
* feat(opengateway): add Macaron V1 Tall to the gateway catalog

Served by opengateway via direct Novita (model is not on OpenRouter).
Free launch window; the gateway delists it 2026-08-10. Adds the model
and brand descriptors and regenerates integration artifacts.

* test(opengateway): add Macaron regression coverage + picker expectation

Adds macaron.test.ts (descriptor capabilities/limits, gateway catalog
apiName/modelDescriptorId wiring, runtime limits — tencent.test.ts
pattern) and includes mindai/macaron-v1-tall in the /model picker's
expected opengateway option list.

---------

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-07-30 11:54:38 +08:00
JATMNandGitHub ae634cdef3 refactor(openai-shim): extract stream lifecycle and response dispatch (#2009)
* refactor(openai-shim): extract client dispatch

* fix(openai-shim): preserve max reasoning effort dispatch

* test(openai-shim): cover injected Codex dispatch
2026-07-29 20:47:28 +08:00
56a920196d feat(web): replace favicon/logo with Ember Block O brand mark (#2065)
The site icon was still the 2026-06 terminal-face + git-fork circuit
mark, predating the ember identity the product now leads with (the
ANSI-Shadow startup logo and the orange pixel wordmark in the README).

Replace it with the Ember Block O: the startup screen's figlet "O"
letterform re-plotted as pure SVG rects — five ember gradient bands
(#ffb15f → #be5008, the exact stops from StartupScreen.palettes.ts)
with the wordmark's thin offset outline shadow, on a dark rounded tile.
Reads as a crisp orange O at 16px and matches CLI, README, and site.

- openclaude-logo.svg: new mark (same filename, Head.astro untouched)
- openclaude.png: 512px transparent-corner render (PNG favicon and the
  nav/footer images, which already reference this path)
- og/{default,docs,commands,buddy}.png: all four social cards
  regenerated with the new mark; layout, copy, and grid unchanged

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-07-29 20:45:24 +08:00
c2030bbb2b fix(web): make web/ build standalone — stop importing the repo-root p… (#2061)
* fix(web): make web/ build standalone — stop importing the repo-root package.json

vercel --prod deploys only the web/ directory, so site.ts importing
../../../package.json (and verify-dist.ts reading it) broke every Vercel
build with ts(2307) while local builds passed.

- SITE.version now derives from latestVersion, the newest entry in
  src/data/releases.ts — committed data inside web/, so builds are
  deterministic and need nothing outside the directory
- verify-dist gains a best-effort npm freshness guard: fails the build
  only when registry.npmjs.org reports a newer @gitlawb/openclaude than
  releases.ts; unreachable registry or malformed responses skip the
  check, and site-ahead-of-npm is allowed for release PRs
- verify-dist.test.ts covers the guard via injected fetch (no network)

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

* fix(web): reject leading-zero semver from the npm registry

Number() would normalize a malformed '01.2.3' to 1.2.3; require strict
semver components so malformed registry values skip the freshness check
instead of being silently coerced.

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

---------

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-07-29 13:12:35 +08:00
0f76b5490b feat(web): v0.26 refresh — buddy page, changelog, partners, provider … (#2060)
* feat(web): v0.26 refresh — buddy page, changelog, partners, provider catalog

- Single-source the site version from the root package.json (never stale again)
- New /buddy/ page: all 7 hero sprites rendered as animated SVGs generated
  from src/buddy/pixelSprites.ts, attack descriptions, commands, hatch lore,
  plus a dedicated 1200x630 OG image composed from the real sprites
- New /changelog/ page: curated release highlights 0.19 -> 0.26 from a typed
  releases.ts data file
- Landing: buddy teaser section, partners strip (GitLawb, Bankr, Atomic Chat,
  Xiaomi MiMo, Atlas Cloud, AI/ML API, Novita AI) with self-hosted logos,
  community links, refreshed provider strip, node >= 22 fix
- Providers docs rebuilt as grouped catalog (39 providers: subscriptions,
  gateways, vendors, local, custom) incl. xAI OAuth, AI/ML API, Cloudflare
  Workers AI, NVIDIA NIM, Kimi K3, GPT-5.6, Opengateway free models
- Data refresh vs v0.26.0 source: 16 new slash commands, pdf skill, new CLI
  flags + 10 subcommands, modelLimits/providerFallbackChain/agentRouting
  settings, corrected env vars (GEMINI_API_KEY, OPENGATEWAY_API_KEY, ...)
- Nav/footer/docs sidebar link the new pages; JSON-LD breadcrumbs on both

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

* fix(web): address CodeRabbit review — flag description + dist verification

- Correct --disable-slash-commands description: the flag empties the entire
  slash-command list (REPL.tsx filters all commands), not just skills; the
  upstream help string "Disable all skills" is the misleading one
- Add scripts/verify-dist.ts, wired into `bun run build` (so the existing CI
  web job runs it): asserts SITE.version matches the root package.json in the
  rendered pages, nav exposes /buddy/ and /changelog/, every release renders
  with its GitHub URL, every hero renders with its sprite asset, partner and
  community links render on the landing page, and the sitemap covers the new
  routes

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

* fix(web): harden verify-dist per review — empty-page guard, rendered-nav check, tests

- page() now records a failure for a present-but-empty file, so '' is only
  ever returned alongside a recorded failure and skipped assertions can no
  longer mask a blank page
- assert the rendered docs sidebar (dist/docs/) links every docsNav route,
  not just the source data array and the landing nav
- extract pure verifyDist(dist) and add 9 fixture-based bun tests covering
  missing/empty pages, lost sidebar links, missing sprites, stale partner
  links, missing release URLs, and sitemap regressions; discovered by the
  root `bun test` run in CI, no workflow changes needed

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

* test(web): derive the missing-sprite fixture from heroes data

Hard-coding robinhood.svg would make the test throw during fixture mutation
if that hero were renamed, instead of exercising verifyDist().

Co-Authored-By: OpenClaude <openclaude@gitlawb.com>

---------

Co-authored-by: OpenClaude <openclaude@gitlawb.com>
2026-07-29 12:32:47 +08:00
10a9190bea fix(ui): keep SpinnerModeGlyph visible inside status parens (#2047)
* fix(ui): keep SpinnerModeGlyph visible inside status parens

Always render the ↑/↓ mode glyph for leader spins so early requesting
and thinking-only phases are not blank, and place it as the first
status part inside the parentheses next to other activity cues.

Closes #2033

* fix(ui): preserve narrow-terminal thinking with mode glyph

Restore a second-chance width gate for leader thinking-only status under
the new inside-parens glyph layout, and suppress the mode glyph when the
row cannot fit minimal status chrome.

* fix(ui): prefer bare thinking over glyph-only on narrow rows

When leader thinking-only cannot fit glyph+thinking chrome, fall back to
bare (thinking) instead of empty mode-glyph status. Tighten glyph residual
budget to account for GlimmerMessage trailing space.

* fix(ui): budget glimmer space in bare thinking fallback

Bare leader thinking-only residual must reserve the GlimmerMessage
trailing space so equality-width terminals do not overflow by one column.

* fix(ui): nest teammate bare thinking under reduced motion

Apply the bareThinkingOnly nested (thinking) wrap in both shimmer and
dimColor branches so teammate thinking-only status keeps parentheses
when reduced motion disables the shimmer arm.

* test(ui): assert exact SpinnerAnimationRow status rows

Fix TS1355 from invalid null as const in baseProps and replace partial
toContain/regex checks with full ANSI-stripped row equality for the
glyph placement regressions CodeRabbit requested.

* fix(ui): prefer status content over empty mode-glyph chrome

When reserving the SpinnerModeGlyph would drop tokens/timer from the
status row, drop the glyph instead. Keep thinking full-chrome recovery,
default unknown modes to down-arrow, and tighten exact-row tests for
typecheck plus CodeRabbit feedback.

* fix(ui): preserve spinner tokens when glyph crowds status

* fix(ui): suppress empty glyph chrome and preserve token recovery

Skip glyph-only status when numeric thinkingStatus cannot fit on narrow terminals, and refuse glyph-free recovery that would swap visible tokens for a timer-only layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ui): tighten glyph recovery for duration and timer bands

Exclude numeric post-thinking duration from glyph-only status (including requesting), prefer streaming tokens over duration when both cannot fit, and recover timer+token rows at the col-30 boundary.

* fix(ui): harden SpinnerAnimationRow glyph recovery priorities

Prefer tokens over timer/duration on mid-narrow rows, recover exact-fit
token columns, keep full effort text when glyph chrome fits, and prefer
active thinking over timer-only status. Tests now use production
Thinking… message width and frozen-clock exact row assertions.

* fix(ui): address CodeRabbit SpinnerAnimationRow recovery nits

Document token-over-thinking tie-break and > vs >= bare-pass split, drop
redundant suppressModeGlyph assignments on token-only fallbacks, and make
exact-row tests use PROD_MESSAGE plus an explicit padded mode-glyph helper.

* fix(ui): drop overflowing spinner suffix for tokens/thinking

Recover mid-narrow status when stop-hook/tool suffixes overflow bare
chrome, prefer live tokens over duration after the drop, and cover long
production verb column bands plus suffix regressions.

* fix(ui): recover status after SpinnerAnimationRow suffix overflow

Re-gate tokens when thinkingStatus is null after dropping a crowding
suffix, restore teammate nested thinking on the same path, and drop the
mode glyph before truncating a suffix that still fits under bare parens.

* fix(ui): complete SpinnerAnimationRow suffix and glyph recovery

Restore timer symmetrically after suffix drop, drop crowding suffixes
when preferTokens would overflow, keep already-visible thinking when
tokens unlock, and prefer tokens over a bare-fitting suffix that cannot
share the row.

* fix(ui): harden SpinnerAnimationRow recovery against wrap cliffs

Budget timer co-restore against all visible parts, re-gate tokens onto
timer-only rows after tokens-over-suffix, prefer thinking over a crowding
bare-fit suffix, and restore the mode glyph only after a suffix-keep
cascade.

* fix(ui): close SpinnerAnimationRow mid-narrow recovery cliffs

Prefer tokens over thinking when they cannot share bare chrome, tighten
thinking-over-suffix exact-fit to avoid a one-column suffix cliff, and
restore the mode glyph whenever recovered leader content fits.

* test(ui): cover SpinnerAnimationRow cliff and glyph-restore cases

* fix(ui): restore tokens beside thinking after suffix recovery

* fix(ui): close SpinnerAnimationRow suffix recovery cliffs

Budget timer and rendered thinking width in suffix-fit predicates so
widening does not drop the elapsed timer or streaming tokens. Co-restore
teammate tokens when thinking crowds timer-only rows, clear the mode glyph
when thinking+token recovery cannot fit glyph chrome, and budget full
effort text before keeping a stop-hook suffix.

* refactor(ui): collapse redundant SpinnerAnimationRow suffix-fit branches

Rely on the combined all-visible suffix budget check instead of
duplicate tokens-only paths. Keeps timer+thinking and no-token
thinking fallbacks unchanged.

* fix(ui): re-gate recovered spinner glyph

* test(ui): tighten spinner layout coverage

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-29 09:40:32 +08:00
JATMNandGitHub 2fe1e1b148 refactor(openai-shim): extract generic and Gemini stream conversion (#2008)
* test(openai-shim): anchor JSON fallback ownership

* refactor(openai-shim): extract stream conversion

* fix(openai-shim): preserve stream completion semantics

* fix(openai-shim): finalize incomplete streams

* fix(openai-shim): reject incomplete tool streams

* test(openai-shim): cover split raw tool text

* fix(openai-shim): reject incomplete fallback tools
2026-07-29 09:39:48 +08:00
0xfandomandGitHub 3925f2791c feat(auth): opt-in loopback proxy hosts that keep subscription (OAuth) auth (#2050)
* feat(auth): opt-in loopback proxy hosts that keep OAuth first-party

Pointing ANTHROPIC_BASE_URL at any host other than api.anthropic.com switches
the client to API-key mode, dropping a signed-in subscription session. That
blocks running the CLI through a local transparent proxy (compression,
inspection, caching) that forwards auth headers to Anthropic unchanged.

Add ANTHROPIC_FIRST_PARTY_PROXY_HOSTS: a comma-separated host[:port] allowlist
that extends first-party detection. It is honored only when the base URL points
at a loopback host, and only loopback entries are considered -- both checks are
redundant by design so a misconfigured non-loopback entry can never widen
first-party status to an off-machine host. Default behavior is unchanged.

Closes #2016

* docs(auth): document ANTHROPIC_FIRST_PARTY_PROXY_HOSTS

* fix(auth): harden loopback proxy allowlist matching

Normalize the base URL port to its scheme default (80/443) before
comparing an explicit allowlist port, so a `127.0.0.1:80` entry matches
`http://127.0.0.1`. Reject embedded credentials and non-http(s) schemes
up front so an OAuth session is never attached to a URL carrying userinfo
or a non-proxy scheme.
2026-07-28 22:39:41 +08:00