Skip to main content
What changed across the Aethis platform. Each entry is derived from the source CHANGELOG.md of the package it belongs to and is labelled with that package. Subscribe via the RSS feed for this page.
aethis-cli
2026-09-27
  • ci: the post-publish unstick step queries every downstream repository at its current path. Two entries still named a pre-transfer path. They worked through redirects, but a new repository created at the old path would have been queried instead and returned no pull requests. No change to the CLI itself.
aethis-cli
2026-09-27
  • ci: the post-publish unstick step now sees repositories that have been renamed or transferred. It listed each downstream repository’s open pull requests with a search filter, and the search index does not follow a repository rename or transfer: a moved repository returned no results and a success exit code, so it was skipped without a warning. The step now lists open pull requests without the search filter and matches the aethis-needs marker in each body, which it already did. A repository that cannot be listed now produces a workflow warning, and so does one that hits the listing limit of 100 open pull requests. No change to the CLI itself.
aethis-mcp
2026-09-25
Authoring safeguards are now visible in tool output. Both are warn-only and never block a generation or a publish.
  • aethis_generate_and_test and aethis_refine render the source questions authoring raised: places where the source text conflicts with itself or can be read more than one way. Each question shows the quoted clauses with their citation keys, the candidate readings, and the provisional reading the ruleset encodes. aethis_publish renders them too.
  • aethis_publish renders the source check: a warning when a cited document is not the text the ruleset was built from (mismatch), when a digest is unavailable (unverifiable), or when the ruleset predates input recording (no_authoring_inputs_recorded).
  • Fix: the generate-and-test path kept only the ruleset id and test result from the final generation status, so fields carried there (source questions and the run’s question counts) never reached the tool output. They are now preserved.
  • All question and check text comes from uploaded sources and model output, so it is returned inside the <api_response> untrusted-content fence.
  • Security hardening: the untrusted-content fence is harder to break out of. It previously neutralised only the exact closing tag </api_response>. Now any occurrence of the fence name in returned text is neutralised, whatever surrounds it, so no spacing, case, lookalike bracket or slash, encoding, or forged opener such as <api_response label="system"> can close the fence or open a new one. Fence labels are restricted to identifier characters, so returned data can no longer reach a fence label’s attribute unescaped.
  • aethis_next_question no longer prints a note’s metadata.type as bare [type] text before the note. That value is returned data, so it now appears only in the note’s fence label.
aethis-cli
2026-09-25
  • feat: print the source check and source questions. aethis publish, aethis rulesets promote-to-live, aethis generate --poll and aethis status now print two warn-only fields when the engine returns them.
    • source_check compares the bytes each citation resolves to at publish with the bytes the ruleset was generated from. The CLI prints it when the status is warnings or error: each mismatch with both digests, each unverifiable citation, and no_authoring_inputs_recorded. An ok or not_run check prints nothing.
    • source_questions lists conflicting, ambiguous or missing source text that authoring raised rather than resolving silently. The CLI prints a count and, for each question, its kind, the quoted clauses, the provisional reading the ruleset uses, the affected criteria and, for a refine, the ruleset it was inherited from.
    • Responses without these fields print exactly what they printed before.
    • Question and warning text comes from uploaded sources and model output. It is never interpreted as terminal markup, and terminal control characters in it (escape sequences, C1 controls, bidirectional overrides, Unicode line separators, lone surrogates) are shown as visible escapes such as \x1b; embedded line breaks become spaces.
aethis-mcp
2026-09-24
Security: provider credentials leave the process only when the user configured them for Aethis. Upgrade recommended.
  • Breaking: an Anthropic key is read from the environment only via the new AETHIS_ANTHROPIC_KEY_ENV server setting, in which the user names the variable holding the key. An anthropic_key_env value supplied in a tool call is refused unless it equals that configured name, so a host model can no longer opt a user’s ANTHROPIC_API_KEY (or any other variable) in on its own. The missing-key error now tells the user how to configure a key instead of suggesting an argument for the model to retry with.
  • Breaking: the retired openai_key argument is refused. Previously it was sent in the Anthropic key header.
  • Any key value that is not Anthropic-shaped (sk-ant-…) is refused locally and never sent.
  • Key-shaped text (sk-ant-…, sk-proj-…, provider-masked echoes) in any upstream response or error is masked before it reaches the MCP client.
  • If you previously exported a provider key in your MCP host’s environment and used the authoring tools, rotate that key.
aethis-cli
2026-09-24
  • chore(examples): retire the bundled spacecraft example in favour of the maintained one. examples/spacecraft-crew-rules/ is removed: its copy of the Spacecraft Crew Certification Act had drifted from the canonical text. The README now points at the maintained example in Aethis-ai/aethis-examples.
  • test(e2e): the spacecraft authoring e2e really runs. It previously pointed at a file path that did not exist and skipped silently. It now fetches the Act, the scenarios and the guidance hints from the maintained example at a pinned commit, verifies the Act’s digest, and fails (never skips) when a fetch fails or the digest does not match. Its own hard-coded guidance and test cases are removed; every maintained scenario is checked via decide. Pinned to aethis-examples 84cad29 (v0.2.7), whose example sources/ directory holds only the canonical Act and its citation manifest.
  • ci(authoring-e2e-weekly): the lane revokes its key after a real run. The revoke step reused the mint step’s Clerk session token, which expires in about a minute. That went unnoticed only because the test used to skip instantly. It now signs in afresh before sweeping the lane’s keys. The lane also refuses to run pytest without a minted key, so a missing key can no longer turn the authoring tests into skips and the lane green.
  • test: the pinned-example check runs on every PR. A credential-free test fetches the pinned example and verifies the Act’s digest in normal CI, so a broken pin is caught per-PR rather than weekly. The redundant 80% pass-rate assertion is removed; the strict per-scenario check covers it.
aethis-cli
2026-09-15
  • fix(generate): preserve structured authored field notes on generation pins. fields.yaml notes, including their opaque metadata, now reach the existing engine note contract. The CLI checks advertised engine support before upload and rejects known unsupported engines.
aethis-mcp
2026-09-14
  • Add aethis_set_tests(project_id, test_cases) for destructive replacement of one existing project’s complete reviewed 1–100-case suite. It verifies the target OpenAPI replacement capability before writing, preserves project sources, fields and guidance, and never creates another project.
  • Replacement POSTs are sent once. Interrupted responses report an unknown outcome and require inspection before another approved replacement.
  • Preserve immutable publication receipt metadata (published_version_id, content_digest, and published_version_label) in aethis_publish output.
aethis-mcp
2026-09-14
  • Resolve the selected CLI profile’s API key and endpoint together, honoring non-secret AETHIS_PROFILE and absolute XDG_CONFIG_HOME references. Profile credentials outrank stale legacy keychain entries; flat legacy files remain supported.
  • Preserve the default endpoint for standalone environment keys, including when the implicitly active CLI profile is anonymous. Explicit named-profile overrides retain the selected endpoint.
  • Refuse invalid selected profiles and unsafe files visibly without logging secrets. Preserve explicit environment overrides and unsigned anonymous setup.
  • Keep startup identity paired until restart; late login refuses an endpoint change before sending a request.
aethis-cli
2026-09-14
  • feat(mcp): secure selected-profile setup for Codex and existing hosts. aethis mcp install now supports codex (and includes it in all) through Codex’s native mcp add/get/remove commands. Every host registration stores only AETHIS_PROFILE and an absolute XDG_CONFIG_HOME, never an Aethis API key or endpoint. A clean install uses the unsigned anonymous profile; saved profiles stay pinned even if the CLI active profile changes later. Conflicting one-off key or endpoint overrides and ambiguous user-managed registrations fail before host configuration is changed. Known legacy CLI registrations migrate safely.
aethis-core
2026-09-07

Added

  • Immutable rulebook releases with two-word labels (#39, workspace epic #1299 P1). Publishing a rulebook now freezes its complete runtime content — every member resolved to exact compiled bytes (plus the published leaf version_id / content_digest), the composition, the field vocabulary overlay, conversational guidance (robot_hints), the scope declaration and the published title — into one immutable RulebookVersion release with its own UUID (release_id), a permanent two-word label (e.g. Amber Heron, allocated from small neutral word lists under a database uniqueness constraint) and a content_identity digest. An identical publication reuses its existing release and label; any change to published content mints a new one. The family’s “latest” is a CAS-protected head on the existing publication-head collection, so a publication is observed complete-old or complete-new, never mixed.
  • The existing operations publish; there is no new endpoint: promote-to-live (always), POST /rulebooks/{id}/activate (always, idempotent — also how a family transferred onto an engine without its releases is published there), and PATCH /rulebooks/{id} / POST /rulebooks/{id}/fields when they supply published content on a published family. Draft families keep pure draft editing. A publication failure is a structured 422 (release_member_unresolved, release_label_allocation_failed, …). Member validation fails before the draft save; a later publication failure keeps the valid draft so the identical content request can resume the idempotent publication.
  • Additive read surface: POST /decide accepts release_id (UUID or exact label, rulebook_id requests only); GET /rulebooks/{id}/schema|explain|graph accept ?release_id=. Responses carry an additive release object (release_id, label, title, version, published_at, content_identity) on /decide, /schema, /explain, /graph, GET /rulebooks[/{id}] (current release) and promote-to-live. A pinned release is evaluated from its frozen content — warm or cold caches — after newer releases supersede it; a release of another rulebook/tenant is 404 with no content; a head naming a release the engine does not hold is a 422 rulebook_release_missing — latest is never substituted. Family calls without release_id resolve the current release when one exists, else the live composition as before (release: null, no backfill).
  • Rulebook /decide reports ruleset_version as v<release.version> when a release was evaluated (previously always unknown); content_identity keeps its composition-key format for existing replay consumers. The route now resolves the rulebook once and hands that composition to the navigator (the route/navigator double-read race recorded on #39 is closed).
  • Behaviour change on published families: any content edit on a family that is active or already has a release (PATCH title/description/ composition/outcome logic/guidance/scope, POST fields) now resolves and VALIDATES every member live under the owner’s visibility before anything is written; a member that is missing, not live, archived, uncompiled or malformed makes the edit fail with 422 release_member_unresolved and the draft is left unchanged. Previously such an edit saved regardless. The draft is then saved and the release published from it; if that last step fails (422 release_publication_failed) the draft is kept and latest lags it by one release until activate (or the next content edit) re-publishes it — latest is always one complete release. PATCH ruleset_refs on a promoted (pinned) family now rewrites the pin map (pinned refs only; a floating reference is 422 composition_pinned) instead of being silently ignored. Promotion resolves every member of the resulting composition before the candidate goes live.
  • Frozen release content is verified against its identity where it is materialised — the cold compile of a release and the schema/explain/graph and overlay projections — and a mismatch is 422 rulebook_release_corrupt. Metadata reads (GET /rulebooks[/{id}]) check the release exists and that its identity agrees with the current head, without hashing member bytes; release_error distinguishes a missing target from a corrupt head tuple. Heads, idempotency tokens and sequence numbers are tenant-scoped, and a release’s sequence number is unique per (tenant, rulebook) by index.
  • Rulebook schema responses include section_names, keyed by section ID. A release freezes each member’s human-authored display name alongside its compiled content, so a pinned application never resolves display text from a newer mutable ruleset document. Releases created before this field expose null for names that were not historically captured.

Fixed

  • GET /rulebooks/{id}/explain returned criteria: [] plus an error for every section with rules: the human-readable renderer was handed a model where it walks a dict (#519). Latent since at least 2026-04-26 on an exposed but uncalled endpoint; found by the release read tests.
aethis-core
2026-09-05

Fixed

  • Rulebook testing publication (#509, #1272). A slug-less candidate can enter testing beside a slug-less current member in the same rulebook slot. Alias-backed replacements retain the existing conflict response.
aethis-core
2026-09-05

Fixed

  • Preserve every existing member when first promoting a legacy rulebook (#505). Resolve legacy references with the decision loader before any writes, pin floating members to the resolved versions, and reject unresolvable or duplicate references before changing publication state.
  • All-green authoring refinement (#507, #1272). Refine mode now asks the authoring model to apply unsatisfied guidance, including metadata-only edits, before its first validation or test call. An already-satisfied request may remain unchanged, while passing goldens alone no longer completes an outstanding guidance edit.
aethis-core
2026-09-05

Fixed

  • Bounded streamed rule authoring (#125, #583, #1272). Long authoring turns now consume the model provider’s complete streaming message, check for cancellation while it arrives, and preserve existing provider/cache usage accounting. A response that reaches its configured output cap cannot dispatch partial tool input; its retained trace is recorded with a typed, safe output-limit failure for a retry.
aethis-mcp
2026-09-03
  • release: verify the Registry’s real response envelope. The official Registry successfully published aethis-mcp@0.17.3 as active/latest, but the final workflow incorrectly read search results from entry.name and entry.version instead of entry.server.name and entry.server.version, so it reported a false-red after publication. Verification now uses a tested parser with a fixture matching the live Registry response shape.
aethis-mcp
2026-09-03
  • release: bind the Registry identity to the case-sensitive GitHub OIDC namespace. Corrects every current MCP release surface from io.github.aethis-ai/aethis-mcp to io.github.Aethis-ai/aethis-mcp, the namespace the Registry grants to this repository. A deterministic test now pins package.json, server.json, the generated tool inventory, and release verification to that canonical identity. aethis-mcp@0.17.2 reached npm but the Registry rejected its lowercase namespace, so it was not a complete dual-registry release.
aethis-mcp
2026-09-03
  • release: satisfy the official MCP Registry metadata contract. Shortens the Registry-facing server description to its 100-character limit and adds a deterministic test for that constraint. aethis-mcp@0.17.1 was published successfully to npm, but the Registry rejected its overlong description during metadata validation, so it was not a complete dual-registry release.
aethis-mcp
2026-09-03
  • release: publish the immutable candidate as a local tarball. The npm publish step now prefixes the downloaded release-artefacts/... tarball with ./, so npm resolves it as a filesystem package rather than a registry package spec. The v0.17.0 workflow failed before publication; no aethis-mcp@0.17.0 package was published.
aethis-core
2026-09-03

Fixed

  • Cloud Run autoscaling verification. Public candidate gates now accept Cloud Run’s valid revision views where the service maximum is omitted or mirrored exactly, while continuing to prohibit revision-level minimums that would keep retired revisions warm.
aethis-core
2026-09-03

Fixed

  • Exact Cloud Run build identity. Deployment gates now preserve BuildKit provenance indexes while verifying Cloud Run’s resolved Linux/AMD64 runtime manifest is the unique runnable child of the immutable build image. The same fail-closed identity check protects staging and public production candidates.
aethis-core
2026-09-03

Fixed

  • Recoverable authoring generation jobs (#481, #434, #420). Generation workers now hold fenced leases with heartbeats and absolute deadlines; abandoned jobs fail explicitly and release their project ownership for an operator-driven retry. Status and cancellation responses expose typed, additive lifecycle telemetry without retaining customer provider keys. Recovery scans are bounded, and operators can pause generation admission while draining old unfenced workers during rollout. Staff job detail uses the same server-authoritative contract version and retry-readiness fields; newer lifecycle records are reported as unsupported and never interpreted as legacy recovery candidates.
  • Safe provider failure details for authoring (#481, #463). Authentication, capacity, rate-limit, request and availability failures are classified into stable reason codes while provider response bodies remain server-side.
  • Date inputs in generated tests (#481). Strict ISO calendar dates are normalised at the navigator boundary, matching the public decision route, and invalid generated test values fail as contained validation errors.
  • Durable public authoring runtime (#481, #485). The public Cloud Run deployment now keeps one instance with CPU available outside requests, limits request concurrency and process-local generation concurrency to one, retains at most one additional local waiter, and declares its CPU, memory, timeout and scaling bounds explicitly. Structured capacity, reaper and admission-pause signals have reviewed alert thresholds. Public staging and production revisions are tagged no-traffic candidates; guarded commands require exact revisions, image digests and database boundaries for drain, legacy recovery, promotion and fence-compatible rollback. Legacy recovery defaults off, and the existing 2 GiB production memory mitigation is preserved independently from staging’s 1 GiB baseline.

Added

  • Compact rulebook choice context for conversational consumers (#477). POST /decide accepts opt-in include_choice_context: true and returns choice_context: complete field-to-section membership plus each authored OR-group’s alternatives, labels, referenced fields and answer-aware open / closed / satisfied status. This avoids constructing and walking the full dependency graph merely to offer an applicant a choice. The existing graph_overlay remains available for visualisation and debugging. timing.total_ms and decision-log latency now include optional response assembly work, so graph/choice overhead is no longer hidden after the core evaluation boundary.
  • Presence-polarity gate at both go-live boundaries (#472; checker from #448; hardened by the PR #475 adversarial review). POST /projects/{id}/publish and POST /rulebooks/{id}/rulesets/{name}/promote-to-live refuse a ruleset whose DEFAULT_TO / IS_UNSET is more generous when the field is never asked than when it is answered, with a structured 422 non_conservative_presence_op carrying field, criterion and solver witness. The outcome-level check runs two variants: each presence field freed singly, and all presence fields freed together — the latter catches OR-composed generous defaults that decide a bundle ELIGIBLE with zero questions asked (review CONFIRMED-1). Stated boundary: proper subsets of silent presence fields between those two extremes are not enumerated. An uncompilable stored bundle fails closed as a violation, never a 500. An internal key may pass force_unsafe: true (now also on PromoteRulesetRequest) to proceed — audit-logged (publish_force_bypass_presence_polarity / promote_force_bypass_presence_polarity). Per the owner decision of 2026-08-23, a republish of an existing reserved-namespace (aethis/*) target — a prior holder of the slug (including the bundle’s own stored slug on a slug-less republish) or a live rulebook member with the same ruleset_name — is ADVISORY instead of blocking: the response carries presence_polarity_advisories and the advisory is audit-logged too (publish_presence_polarity_advisory / promote_presence_polarity_advisory), because aethis/construction-all-risks’s two accepted defaults are load-bearing for published benchmark artefacts (Aethis-ai/aethis-examples-internal#38). A brand-new first-party slug or member name binds like anyone else. The advisory/binding decision is one named function (presence_polarity_gate_mode), so retiring the carve-out is a one-line, cited change.
aethis-sdk-python
2026-09-02

Added

  • Typed recovery support for asynchronous ruleset generation (aethis-core#481): Aethis.get_generation_status() / AsyncAethis.get_generation_status() now return GenerationStatusResponse, including structured job progress, worker-heartbeat, lease/deadline, test, and terminal-failure telemetry.
  • Aethis.cancel_generation() / AsyncAethis.cancel_generation() explicitly request cancellation of an observed generation job by project and job id, preventing a delayed request from cancelling a successor job, and return GenerationCancellationResponse. Cancellation is cooperative: it fences the job and releases the project, while an in-flight provider request stops at its next safe boundary. The SDK never cancels, resumes, retries, or stores credentials on a caller’s behalf. Requires the aethis-core generation- recovery API to be live on the target API; earlier engine versions do not advertise generation_contract_version=1 and are refused before mutation. Status exposes telemetry availability, server-authoritative worker lifecycle, and retry readiness; cancellation distinguishes cancelled from the idempotent already_cancelled outcome.
aethis-mcp
2026-09-02
  • feat: generation status and cancellation. Adds aethis_generation_status to inspect a project’s current or most recent authoring job, and aethis_cancel_generation to abandon a target-bound observed job (project id plus matching job/confirmation ids) and release project ownership. Worker shutdown may be cooperative rather than immediate. Status is read-only; cancellation is explicitly annotated as a destructive API-key mutation, so MCP hosts can gate it for approval. Matching ids prevent retargeting but do not prove human consent; agents must obtain a fresh explicit reply before calling. Both responses and diagnostics remain fenced as untrusted API data. Status carries telemetry availability, server-authoritative worker lifecycle, and retry readiness; cancellation preserves the idempotent cancelled / already_cancelled outcome.
  • safer authoring timeout guidance. A timed-out aethis_generate_and_test now tells agents to inspect status before retrying and to cancel only when the caller asks to stop the run. This release requires aethis-core’s project status and generation-cancel endpoints to be live.
aethis-cli
2026-09-02
  • feat: explicitly recover an authoring project from an abandoned generation. aethis cancel [-p PROJECT] first observes and displays the exact active job, then calls the engine’s job-bound cancellation endpoint after a confirmation (--yes for automation), marks that job failed, and releases its ownership of the project. It reports the engine’s honest limitation: an already-running worker may continue even though a new run can now be admitted. The idempotent response distinguishes cancelled from already_cancelled. Polling never cancels automatically.
  • feat: generation status shows live convergence and heartbeat telemetry. aethis status and the live generate progress line consume the engine’s turn count, best test pass rate, tool count, last tool, and seconds_since_progress. Older engines remain compatible: absent telemetry is omitted rather than guessed. A client-side timeout now names the exact status and cancel recovery commands. Status also renders telemetry availability, server-authoritative worker lifecycle, and retry readiness.
aethis-cli
2026-09-02
  • fix(generate): a dropped connection mid-poll no longer strands the run on a stale ruleset id. --poll guards against an interrupted generation leaving .aethis/state.json naming an earlier ruleset, but the guard only covered AethisAPIError. A dropped connection, DNS failure, TLS error or read timeout raises httpx.HTTPError, which unwound past it — so the failure most likely to interrupt a long poll was the one that bypassed the guard, and every later decide, test, explain and fields pull silently answered from the wrong ruleset. The poll now absorbs a short connection blip and finishes normally; where the API is genuinely unreachable it names the stale id and prints the command that recovers the real one.
  • fix(generate): the stale-pointer guard can no longer fire on a run that succeeded. The pinned-vs-produced field diff runs after the new id is recorded, so a connection lost during it made the CLI announce “Done! Ruleset: X” and then insist X was from an earlier generation — a false claim about the exact fact the guard exists to keep honest. The diff now honours its own “never fails the command” contract for transport errors, and the guard never describes a pointer this run recorded as stale.
  • fix(generate): a blip while waiting for the new ruleset id no longer discards it. After the engine reports success the CLI re-polls for the id. A transport failure there stranded a generation the CLI already knew had succeeded; those attempts are now individually tolerant.
  • fix(generate): a persistent outage is reported as unreachable, not as a timeout. The retry is bounded by consecutive failures, so a dead connection surfaces as one rather than being absorbed to the deadline and reported as a slow job. A flapping link is deliberately not aborted — it is working, and killing it one poll short of success would deliver the very failure above — so a poll that suffered dropped connections says so when it does time out.
  • fix(generate): a failed publish no longer reads as a failed generation, and says why. A transport error on the auto-publish reported only “Could not reach the Aethis API”, so a successful generation looked like a failure and got re-run. It now reports the ruleset, names the reason the publish did not run, and names aethis publish as the step to re-run — previously the API-error half of this path gave a green tick with no reason at all.
  • fix(generate): a success that never surfaces an id says so. That ending also leaves the recorded pointer naming an earlier generation, and was silent.
  • fix(generate): an unreadable value space is reported, not raised. The field diff verifies value-space-pinned fields against the registry and formatted the failure as if it were always an API error, so a dropped connection there raised AttributeError — an unhandled traceback on an otherwise successful generation.
  • fix: an unreachable API cannot crash the error reporter. httpx exposes .request as a property that raises when unbound, so the defensive getattr(e, "request", None) in the transport-error path could raise RuntimeError — a traceback about the error handler in place of the one actionable line. The rendering now has a single home shared by the top-level boundary and the commands that must clean up before exiting. generate and refine name the host they were actually configured with rather than the public default; other commands still report via the top-level boundary, which reads AETHIS_BASE_URL.
aethis-core
2026-08-23
Measured against the deployed 0.57.0: additive on the public wire (owner probes of an exact private draft via /decide, two opaque per-field metadata keys on both schema routes, durable authoring-source identity) plus one navigation correction that changes which question /decide asks — never the verdict — on rulesets with alternative Boolean branches.

Fixed

  • A Boolean route is no longer flattened into its field union — dead branches stop driving questions (aethis-core#465, #466; defect shape DS-63; workspace epic ws#1067). For a criterion such as (kind = ordinary AND taught) OR (kind = certificate AND aquals AND elps) the compiler always produced the right condition, but the navigator also kept the flat union {kind, taught, aquals, elps}, and twelve consumers — find_next_input_field, relevant_unanswered_fields, _check_requirement_status, get_best_requirement, get_optimal_path_to_eligibility, _add_input_field_constraints and its tie-breaks, the layered explanation’s unused_facts, routing.py, the scenario evaluator, build_pydantic_model, detect_review_triggers — treated that union as though every member were conjunctively required. With kind = certificate the engine could still ask taught, report it missing, cost it on the optimal path, or hold the requirement at unknown. Each EligibilityRequirement now carries one requirement-scoped, linear NNF semantic plan built from its unresolved AST (NOT pushed to literals, IMPLIES(a,b) lowered to OR(NOT a, b), no DNF materialisation, no route-count fallback). Decision/status evaluate it under _resolve_presence_ops; askability, selection and relevance evaluate the same plan under _relax_presence_ops, so an unanswered DEFAULT_TO / IS_UNSET fact that could still change the outcome stays askable. Duplicate criterion ids across composed sections keep requirement-scoped identity and collision-free solver selectors. Non-applicant fields remain selectable at a high finite cost so an applicant-completable branch wins when one exists. The flat input_fields union survives as provenance/schema metadata only; if the engine cannot attribute the authored structure it fails loudly rather than asking the union. Verdicts are unchanged on every existing rulebook (the DSL parser and truth evaluation were never shown to be wrong; read-only scan of all 257 Rules_Staging bundles: 0 flat bundles with authored outcome_logic). Visible correction: status_report["path"] may now name the actually satisfied requirement where the old flat completion proxy returned None. Reproducer: the private English-language draft asked eng.degree_taught_in_english on the Australian postgraduate-certificate route (6/7 question-order probes on main; 7/7 after). A narrow missing_fields-only over-report on two flat-bundle shapes is tracked in #467; decision, next question and optimal path are correct there too.
  • Explicit field pins are authoritative through the single post-parse normalisation seam (#469, #470): authored label, question, weight, owner, injection, recovery and UI metadata on a pin survive fresh generation and refinement alike, absent-vs-explicit-null is preserved at ingestion, and the same path serves inline and named value-space fields.
  • Field pins are enforced inside the authoring TDD loop, not only at persistence (#460, #461): compile_and_test applies normalise_parsed_fields before compiling, so a missing/unparseable pin or invalid space expansion reaches the model during iteration instead of killing the job at the end.
  • Authoring normalisation failures keep their reason code, affected fields and parse warnings in the job error detail (#458, #459), and the authoring preflight’s exact token count is emitted through the deployed application logger (#456, #457) — observable on staging without logging prompt or source text.

Added

  • Presence-operator polarity checker — report-only (#448; refs tda-server#1351). find_non_conservative_presence_ops (per criterion) and find_outcome_polarity_violations (per composed outcome) detect a DEFAULT_TO / IS_UNSET whose never-asked reading is more favourable to the applicant than answering it would be — a satisfying model of C[all defaulted] AND NOT C[f free] is a concrete world in which declining to ask f passes while some real answer fails. Both readings come from the navigator’s own _resolve_presence_ops / _refresh_conditions_from_answers, so the check cannot drift from the semantics it certifies; it fails closed. Measured: form-an-english-language-r2 has no presence ops; form-an-life-uk-r2 and aethis/consumer-credit-prequalification are clean; aethis/construction-all-risks has 2 violations (verdict flips with ask order, witness-driven through the shipped navigator). Nothing calls it yet — binding it would break that published ruleset on contact (adaptive-gating.md rule 5); the measurement above is what the gate decision now has.
  • Authenticated ruleset owners can evaluate an exact private draft through /decide without publishing or activating it (aethis-core#464; workspace epic ws#1067). The fallback accepts only the immutable compiled-ruleset id and the owning tenant: draft slugs, project aliases, anonymous callers and other tenants still receive 404, and the request remains read-only. This gives authoring acceptance harnesses a supported way to test question order on the exact generated artefact before it can become live.
  • Rule-authoring inputs now have durable identity and exact preflight diagnostics (aethis-core#452; workspace epic ws#1081). Identical concurrent uploads within one tenant/project/filename reuse one stored source identity and report new versus reused counts; inactive matches and changed active filenames return actionable PATCH-lifecycle conflicts. New generated citations use immutable source_id#item_id keys, and source hard deletion is disabled in favour of soft supersession/reference-only status. Both generation routes assemble one immutable prompt, make one authoritative provider token-count call, and execute that same prompt object. Rejections report the exact total plus deterministic UTF-8 byte/character sizes by component without claiming per-component token attribution. Empty eligible corpora and unresolved generated citations fail before ruleset persistence; concurrent requests are admitted through the existing project state.

Changed

  • Internal production engine scales to zero; staging stays warm (#462). The $_ENVIRONMENT branch in cloudbuild.yaml now sets min-instances per environment; part of the August cost response (ws#503).
aethis-cli
2026-08-22
  • fix(generate): authored field behaviour is deterministic. Generation uploads now carry explicitly authored labels, questions, ownership, injection source/phase, ordering weight, recovery capabilities, and UI hints on the field pin. The engine’s normalisation seam can therefore preserve the authoring contract instead of accepting model-invented replacements.
  • fix(generate): an omitted question on a non-applicant field is explicit. When a field declares a caseworker, system, or derived owner and has no question, the CLI sends question: null; this clears any question invented by the generation model. Legacy fields that declare none of the new metadata retain their previous upload shape.
  • fix(generate): capability checks cover behavioural metadata. An older engine that would silently discard any declared property is refused before the field-spec push, just as it already was for enum labels and canonical storage mappings.
aethis-cli
2026-08-20
  • fix(generate): source upload results are truthful. Generation reports the engine’s exact new and reused counts instead of calling every attempted file newly uploaded.
  • fix(generate): edited sources remain regenerable. Before uploading changed content under an existing filename, the CLI supersedes the prior active source, uploads and links its replacement, and reactivates the prior source if the replacement upload fails. This avoids both a hard-blocked edit loop and duplicate active source text in the authoring prompt.
aethis-sdk-python
2026-08-19

Added

  • SchemaField reads the two authored metadata properties the engine now publishes (aethis-core#449; workspace epic aethis-workspace#1067). enum_labels — an optional {member slug: display label} map — and canonical_field — the author’s pairing between an eligibility field key and the consumer’s canonical storage key — are carried on every field of both /rulesets/{id}/schema and /rulebooks/{id}/schema. Previously the SDK’s model declared neither, so Pydantic silently dropped them and an SDK caller could not render a label for a stored enum slug. Both are opaque to the engine and to this SDK: nothing here validates enum_labels against enum_values, or resolves canonical_field. Both are optional and default to None, so a schema from an engine predating the properties parses exactly as before. None and {} are deliberately distinct — {} is an author declaring “no labels”, None is an author declaring nothing — so read them with is None, never truthiness.
aethis-cli
2026-08-19
A field can now say what its members are called, and which stored key its value belongs to. Requires an engine that models both on a field spec — one that does not is refused before the upload, never allowed to accept it and drop them. Projects declaring neither are unaffected.
  • feat(fields): enum_labels: on a field pin. An enum field may carry a per-member map of display wording — enum_labels: {ion_drive: Ion drive} — beside the members themselves. The engine treats the labels as opaque text and publishes them on both the ruleset and the composed-rulebook schema, so every consumer renders the same wording from one authored source instead of each keeping its own copy of it. Optional per member: label the ones whose slug does not read well and leave the rest. Both properties are handled on presence, not emptiness: an explicit enum_labels: {} is a declaration the engine keeps and republishes as {}, so it is transmitted, validated and preserved across a fields pull rather than being discarded as an empty value — and it fires the capability check below like any other declaration.
  • feat(fields): canonical_field: on a field pin. The storage key a field’s value belongs to (canonical_field: spacecraft.propulsion) is authored beside the field and published on the schema, rather than re-derived from the key by every consumer that needs the pairing. Any field type may declare it.
  • feat(fields): both are validated locally before they can reach the engine. enum_labels must be a mapping of member → non-empty text on an enum field, and where the members are declared inline it may not label a member the field does not have — the engine accepts such a label and simply never renders it, which is invisible at every layer afterwards. Where the members are a named value_space they live on the registry, so only the shape is checked. canonical_field must be non-empty text.
  • feat(generate): an engine that cannot keep them is named, not written to. A field-spec property an engine does not model is ignored rather than rejected, so the upload succeeds, the generation runs, and the authored values are gone — the same silent-drop shape the value-space pin was guarded against. The CLI reads the engine’s own published schema and refuses before the push, naming the properties and the engine. An engine whose schema could not be read at all is a different answer and says so: the upload proceeds with a warning to check the published schema afterwards, because unknown is not unsupported.
  • Unchanged for everything else. A project declaring neither key produces the byte-identical upload payload it did before, and never triggers the capability probe — so it keeps working against any engine, of any vintage.
  • feat(rulebooks): set-fields is guarded the same way. aethis rulebooks set-fields posts its fields file as authored, so it already carried both properties — and would equally have had them silently discarded by an engine that does not model them. It now refuses first, asking about the rulebook field model rather than the project one, since an engine can carry the properties on one and not the other. A fields file declaring neither key is untouched and never probes. That command still performs no other validation of its file (#114).
aethis-core
2026-08-19
Tagged 2026-08-19 without a changelog entry; documented here after the fact. Additive on the public wire.

Added

  • next_question says which section(s) the question serves (#446). NextQuestion.sections: List[str] (default [], so older callers see no change) lists every composed section whose requirements read the selected field — a set, not a single value, because the first question of a multi-section rulebook is typically a shared fact. Populated on next_question and on the optimal_path fallback; derived in place from section_to_groups with no extra solve, and fails soft to [] so a narration aid can never cost a decision. Also exposes EligibilityNavigator.sections_for_field(name).
  • Two pieces of authored per-field metadata now ride the published schema (aethis-core#449; workspace epic ws#1067). enum_labels — an optional {member slug: display label} map — and canonical_field — the authored pairing between an eligibility field key and the consumer’s canonical storage key — are carried from authoring to both /rulesets/{id}/schema and /rulebooks/{id}/schema, through ExpectedFieldSpec ingestion, the DSL emit/parse round trip, the compiled FieldDefinition, and the rulebook field vocabulary (where a declaration overrides the member ruleset’s, like question). Both are opaque to the engine: nothing compiles, solves, gates or validates them, and enum_labels is deliberately not checked against enum_values — a referenced field’s members come from the value-space registry and may move independently of its labels. They exist so a consumer can render a human label from the single stored slug instead of keeping a second copy of the vocabulary; a hand-maintained copy of the Form AN enum drifted from the authored source and silently discarded an applicant’s answer, which is what this deletes. The pin is the sole origin. A pin that declares a property sets it; a pin that is silent CLEARS it; a field with no pin cannot carry either. So a property can be removed by authoring, not only added, and a model cannot invent one. For a field shared by several rulesets in a rulebook, the first member in the rulebook’s pinned order that DECLARES a property supplies it — which is not the same as the first member that merely mentions the field. Additive and optional throughout. A build that omits them persists neither key, on both of the populations that store them — the compiled field (they join _CONDITIONAL_FIELD_KEYS) and the rulebook’s own field vocabulary (guarded at both the Pydantic and the BSON layer, because Beanie encodes nested models without calling model_dump(), and every rulebook write is a full-document save). So a field authored without them is byte-identical to today and a rolled-back engine still re-validates it. Both schema routes emit an explicit null so a consumer can tell “no labels authored” from “this engine predates labels”.
aethis-core
2026-08-17
Measured against the deployed 0.49.2 on 2026-08-17, this release is purely additive on the public wire: three new routes, nine models gaining fields, four new schemas, and nothing removed or renamed. No consumer change is required to keep working.

Added

  • Named value spaces (aethis-core#423/#428, #424/#429, #426/#432; workspace epic ws#980). A model may now reference a shared, named vocabulary instead of retyping the same enum on every field. GET/POST /api/v1/public/value-spaces and GET/PUT /api/v1/public/value-spaces/{name} expose the registry; ExpectedFieldSpec gains value_space; expansion happens deterministically at compile time, so a decision is never at the mercy of registry state at solve time. Ships with drift telemetry over the generation record.
  • Engine-level guidance tier (#427/#430). GuidanceListItem and AddGuidanceRequest gain tier, overrides, adherence, principle_key and process_type, with admin routes at /api/v1/admin/guidance/engine. Adherence is reported honestly rather than assumed.
  • A rulebook declares whether its outcome is a COMPLETE determination (#412). complete_determination on CreateRulebookRequest, UpdateRulebookRequest, RulebookResponse and DecideResponse, so a caller can tell “this rulebook has decided everything it governs” from “this is all it can say”.
  • Typed undetermined_reason on /decide (#402, ws#884). The positive terminality signal: consumers no longer have to infer why a case is undetermined from the shape of what came back.
  • pending_non_applicant_input, plus elicitation_owner and recoverable_from on NextQuestion — who may supply a given fact is now part of the wire contract rather than something a caller reconstructs.
  • POST /api/v1/public/projects/{project_id}/generate/cancel and live turn/convergence telemetry during authoring (#403).
  • Replace semantics for POST /tests (#404) — adding a suite no longer silently duplicates it.
  • A production deploy gate (#441, closes #442). A tag previously ran no tests at all: test.yml has no tags: trigger and cloudbuild.yaml has no test step. scripts/preflight_production.py now blocks a tag unless the tree is clean on main, CI is green on that exact commit, staging is verifiably running the code being promoted, and a behavioural smoke passes against it; .github/workflows/post-deploy-verify.yml re-runs that smoke against prod after the roll and opens an issue if the deployed engine stops deciding correctly. The oracle lives in scripts/smoke_expectations.json.

Fixed

  • A group dead only by a presence-default keeps offering the reviving field (#418) — the navigator no longer closes a path the applicant could still open.
  • The final best-source re-evaluation carries the value-space context (#435), so a field resolved against a named space is not re-scored without it.
  • An authenticated owner may read their own draft ruleset’s schema/explain/graph (#436). Ownership, not publication status, governs the read.
  • Deploy-time database guard (#408/#409): DATABASE is derived from _ENVIRONMENT, and a prod deploy against a staging database is refused outright rather than discovered afterwards.
  • RULES_AUTHORING_MODEL has a durable, staging-only home (#405/#406) — --set-env-vars replaces the whole env, so a hand-set value was wiped by the next deploy.
  • Ruff rule set pinned via explicit select (#415); the 0.16 default broadening had turned the daily dependency heartbeat red.

Changed

  • Model catalogue refreshed to the gpt-5.6 family (#401), repointing the mini class like for like.
aethis-cli
2026-08-16
A member section’s field pin is its own contract, not the rulebook’s.
  • fix: rulebook-level fields no longer extend a section’s pinned field spec. The merge pinned every rulebook field onto every member section, demanding fields the section’s rules never author — and the engine historically dropped them silently (every published referees_identity build was missing 3-5 of the rulebook’s cross-section fields). With the engine’s new loud pin-presence gate, such a generation now correctly fails — which surfaced the over-broad pin on the first live canary run of the named-value-spaces epic. The rulebook’s definition still wins on keys the section declares itself (the canonical definition is unchanged); it just never adds keys. The drift report’s comparison universe changes identically, so it now reports on exactly the section’s own contract.
aethis-cli
2026-08-15
Named value spaces reach the wire (aethis-core#424, epic aethis-workspace#980 P3). Requires an engine exposing /api/v1/public/value-spaces for referenced fields; projects with no value_space: pins are unaffected.
  • feat(fields): value_space: on a field pin. fields.yaml enum fields may reference a named, versioned value space (e.g. value_space: form-an/countries) instead of inlining enum_values — the engine expands the members deterministically and the model never authors them. The two are mutually exclusive per field, validated locally; reference-form enums no longer fail the “enum needs enum_values” check.
  • feat(generate): registry sync before spec-set. Locally-authored space files (value_spaces/ or shared/value_spaces/*.yaml, wire-form name/members/provenance) are PUT to the engine registry before the field spec is pushed, carrying the sync-state base version so a stale checkout gets the engine’s 409 (“space moved under you — pull first”) instead of silently regressing the shared vocabulary. Any non-2xx aborts before spec-set naming the missing engine capability — an older engine ignores unknown pin keys and would otherwise silently drop the reference.
  • feat(generate): registry-aware drift report. A referenced field is verified against the registry at the exact version the generation resolved (from the job result), printing <key> ← <space>@vN (M members verified) — never the old “pins no enum_values ⇒ opts out” silence on exactly the migrated fields. A space that advanced between sync and generation is flagged. An unverifiable referenced field is reported, never skipped.
  • feat(fields): pull preserves the migration. A local field declaring value_space: keeps the reference; the server’s expanded members are never written back into fields.yaml.
aethis-cli
2026-08-15
The releases start reaching PyPI again.
  • fix(packaging): explicit package discovery. Every release from v0.30.0 to v0.34.0 was tagged and never published: the publish workflow’s own integrity step creates a top-level evidence/ directory before building, and setuptools flat-layout auto-discovery aborts the build the moment any second top-level directory exists (#94). Discovery is now explicit (include = ["aethis_cli*"]), so a stray directory — the workflow’s or anyone else’s — can no longer poison the build. Proven both ways: with evidence/ present the build failed before this change and succeeds after it. This release supersedes v0.30.0–v0.34.0, none of which reached PyPI; it is the first PyPI release to carry everything since v0.29.0, including --no-publish and the member-set drift report.
aethis-cli
2026-08-15
Generating a ruleset no longer has to activate it.
  • feat: aethis generate --no-publish leaves the ruleset an unpublished draft. A successful aethis generate --poll publishes what it produced, and publishing activates it. For authoring that must leave a draft behind — a ruleset that only ever activates somewhere else, after promotion — there was no way to opt out, and the only workaround was to archive it immediately afterwards, which leaves a real window where it is live. The flag suppresses the publish and nothing else: the poll, its timeout, and the post-generation field diff are untouched, because those are what make the run worth doing.
  • feat: the ending says which one it was. A run that was told not to publish reads differently from one whose publish failed. Both leave a draft and both point at aethis publish, so collapsing them would tell an author who never passed the flag that everything went to plan — precisely when they most need to know it did not. The deliberate ending names the flag; the failure ending is unchanged.
  • Unchanged without the flag. It defaults off, and aethis refine, which shares the same machinery, was not asked to change and does not: the publish still happens on the same call with the same argument.
aethis-cli
2026-08-15
A pinned enum’s members are checked, not just its name — and the check survives the failure it was written to explain.
  • fix: a failed generation now prints the drift report before it exits. The diff existed for the case where the model did not honour the pin, and that was the one case it could not reach: the poll loop exited the moment a job reported failed, several frames below the call that would have printed it. A failed generation said “Generation failed” and nothing about what had actually been produced. The poll now reports how the run ended and the caller decides, so the diagnostic runs first and the exit code is unchanged.
  • fix: the diff is computed against the ruleset this run produced — never the last one recorded. It previously read the ruleset id off .aethis/state.json, which is written only on success, so a run that produced nothing would have been diffed against an earlier generation and the result would have read exactly like this one’s. Where a run names no artefact — which is every failure today, since the engine records one only on its success paths — the CLI says so plainly instead of falling back on anything.
  • fix: the ruleset a successful generation reports is the one that job produced, not the project’s newest. The id was read from latest_ruleset_id, which is a property of the project — so with two generations running against one project it could name another run’s artefact, both in the diff and in the id written to disk for every later command to default to. The job’s own result_ruleset_id is now preferred, with the project-level value kept only as a fallback for engines that record nothing on the job.
  • fix: a generation whose polling breaks off no longer goes quiet about the recorded ruleset. An API error mid-poll — a 500, a dropped connection — exits without a verdict, and the job may well have succeeded regardless. Like a timeout, nothing has been ruled, so the recorded id stands; unlike before, it is named, because that ending was otherwise the remaining way to reach a silent aethis fields pull from an earlier generation.
  • fix: a failed generation no longer leaves a ruleset pointer that reads like its result. The recorded id is what aethis fields pull defaults to, so after a failure the next pull synced from the previous generation without a word, and the fields appeared to be the ones just generated. A failure now clears it — naming it, so it can still be passed with --ruleset-id — and fields pull refuses rather than guessing. A timeout is treated differently on purpose: nothing has been ruled, the job may still land, so the pointer stands and is named as stale instead of discarded.
  • fix: an enum that came back with no members is reported as drift, not success. The member comparison skipped a produced enum whose member set was empty, reading it as “nothing to compare” — so a field whose members had all been dropped printed ✓ Fields: all N pinned field(s) were produced. That is the worst case the report exists to catch, reported as the best one. The empty set is now the loudest result, and a field that came back as something other than an enum reads the same way.
  • fix: a schema that cannot be read is said out loud. The fetch failure was swallowed entirely, so a purged or unreachable draft produced a run with no verdict at all — indistinguishable from a clean one. It is now reported, naming the ruleset, and still never fails the command.
  • fix: aethis generate compares pinned enum member sets, not only field keys. The post-generation drift report asked whether each pinned field was produced. It was — so a schema whose enum had quietly grown a member nobody pinned printed ✓ Fields: all N pinned field(s) were produced. The members were already in hand locally; only the comparison was lossy. It now checks each pinned enum by exact equality in both directions and names the added or dropped members.
  • Why it matters beyond tidiness. A field can be present, correctly typed, and still wrong. Where an enum is used as an escape hatch — a value that keeps a decision at undetermined pending human review — the safety property is that no value the schema allows can turn that into a definitive “no”. One unpinned member breaks it, and no test case can catch that: a test occupying the offending value would itself be the bug. Five consecutive generations of a real ruleset were scored as clean this way.
aethis-cli
2026-08-05
Re-uploading a test suite no longer leaves a second copy of it.
  • fix: aethis generate replaces the project’s test cases instead of adding another copy of them. tests/scenarios.yaml is the authoritative suite and it was uploaded in full on every run, while the API only ever appended — so N authoring cycles left N copies of every case. Nothing errored: duplicates inflate the denominator of every pass rate, so a run reported a total that looked like a result and was partly copies of itself. Two real projects were found holding 185 and 96 cases against files of 42 and 14. The upload now asks the engine to replace, which makes it idempotent.
  • feat: the count of cases the upload overwrote is reported. Replacing is destructive — it removes what was on the project — so aethis generate prints how many cases went and how many arrived (Uploaded 42 test case(s) from scenarios.yaml — 42 replaced) rather than performing it silently.
  • feat: an engine that cannot replace is named, not worked around quietly. Replacing needs an engine that offers it, and one that does not offer it does not reject the request — it ignores the member and appends, so sending it blindly would restore the duplication with nothing to notice. The CLI reads the engine’s own published schema, sends the flag only where it is advertised, and where it is not — or where the schema could not be read at all, which is a different answer and says so — prints that the upload APPENDED, that running again adds another copy, and what to do about it. Nothing is dropped silently in either direction.
aethis-core
2026-08-02

Fixed

  • A rulebook’s field vocabulary now reaches the navigator /decide solves against, not only the /schema projection (aethis-core#396). The rulebook is the surface on which a fact’s ownership is corrected without regenerating a member ruleset. 0.53.0 implemented that correction in exactly one place — the GET /rulebooks/{id}/schema route — and nowhere on the solve, so the two public surfaces disagreed about the same field in the same request window. Measured on the live engine: /schema reported elicitation_owner: caseworker, while /decide proposed that same field to the applicant as next_question, labelled it applicant, and returned pending_non_applicant_input: null. Ownership itself was never broken — the engine honours it exactly whenever the compiled field carries it. The overlay simply never reached the navigator, which is compiled from the members.
    • elicitation_overlay() (public/models/rulebook.py) is now the single definition of what a rulebook’s vocabulary says about a field’s four elicitation axes. /schema, the navigator and /decide’s field registry all read it; none re-implements the merge. /schema’s output is unchanged.
    • The overlay is a LEFT JOIN applied to the compiled navigator, and it can restrict askability but never grant it — so a vocabulary entry added purely to reword a question cannot silently re-open a caseworker fact to the applicant. The effective owner is the union of “the compiled field says non-applicant” and “the rulebook says non-applicant”; an elicitation_owner of applicant, and an empty recoverable_from, contribute nothing. That union is not a stylistic choice — it is the only merge that is safe on data written by 0.53.0, which materialised elicitation_owner: "applicant" and recoverable_from: [] into every stored spec. No later read can distinguish those bytes from a deliberate declaration (workspace defect-shape DS-43), so any “the vocabulary wins where declared” rule downgrades a compiled caseworker field on every legacy row. Under a union those bytes are inert whatever they meant. injection_source and injection_phase were Optional in 0.53.0 too, so for those absence is observable and they use is not None. The cost is that a rulebook cannot re-open a compiled non-applicant field to the applicant. Measured before adopting it: 0 of 352 compiled fields declare a non-applicant owner, so the capability has no current use case, and the asymmetry runs the safe way — a withheld question surfaces as a visible pending_non_applicant_input entry, whereas a wrongly-asked one is the defect this whole contract exists to prevent. It changes who may supply a fact, never whether the fact constrains the solve — the omitted field still decides the case when its value arrives.
    • Corrections work in both directions: a rulebook may also declare a compiled caseworker field applicant-owned, and the engine will then ask it.
  • pending_non_applicant_input no longer degrades silently. The projection swallowed every exception into [], which is indistinguishable on the wire from “nothing outstanding” — so a broken projection reported every waiting case as finished, the exact failure the field exists to prevent. The documented tolerance (a navigator predating the projection) is kept but narrowed to AttributeError; anything else is logged at ERROR with a traceback. Either way the caller is told — and a degraded body is never written to the decision cache, because the header that marks it degraded is a per-response artefact that does not ride the cached payload. Caching one would have served “nothing outstanding” with nothing to contradict it for the full TTL (24h for a pinned rulebook) — the same silent degrade, one layer down. Healthy bodies are cached as before.

Added

  • X-Aethis-Pending-Input response header on /decide. Present only when the outstanding-non-applicant-facts projection could not be computed (unavailable; reason=unsupported|error|malformed). Absent on every healthy response. Consumers that treat pending_non_applicant_input: null as “nothing outstanding” should treat a response carrying this header as “unknown” instead. Mirrors the existing X-Aethis-Records-Omitted header on the catalogue route.

Changed

  • RulebookFieldSpec’s four elicitation axes are now Optional, defaulting to None instead of to "applicant" / []. None means the author did not declare one, and the compiled field’s value stands; it is not a new owner class. GET /rulebooks/{id}/fields therefore returns null for an undeclared axis on documents written from 0.54.0 onward. Rows already in the database were written by the 0.53.0 model and still carry their materialised "applicant" / []; there is no migration, and none is needed, because the overlay’s union rule (above) makes those values inert. GET /schema still emits an explicit "applicant" when neither the vocabulary nor the compiled field declares an owner.
  • GET /rulebooks/{id}/schema changes its answer for some existing stored documents. The 0.53.0 route treated any vocabulary entry as authoritative for all four axes; 0.54.0 treats a stored injection_source: null / injection_phase: null as absence and falls through to the compiled field, and ignores a vocabulary elicitation_owner: "applicant" / recoverable_from: [] entirely. A differential over stored-document shapes finds 350 of 512 stored-row combinations diverge — e.g. a vocabulary entry storing nulls over a compiled definition that declares a source and phase now reports the compiled values rather than None. Every divergence is in one direction: the compiled declaration is no longer silently erased by a materialised default, and no field becomes more askable than 0.53.0 reported it. That direction is asserted as a property test, not just observed. On the currently deployed vocabulary the served bytes are unchanged, which is why this was initially and wrongly recorded as “byte-identical”: 0 of 352 compiled fields declare injection_source / injection_phase / recoverable_from, and all 3 stored vocabulary entries declare elicitation_owner: caseworker, so none of the 350 shapes occurs in production today. That is a fact about the data, not about the code.
  • content_identity for a rulebook decision binds the rulebook’s field vocabulary — but only when the rulebook actually carries corrections. The vocabulary determines what the engine asks and what it reports as pending, so a decision taken under one vocabulary is not replayable against another, and an identity ignoring it claimed a replayability it did not have. A rulebook with no corrections keeps its 0.53.0 identity byte-for-byte (the component is appended only when present, never as a :none sentinel): nothing about how it decides has changed, and consumers compare recorded against replayed identities to decide whether to serve a decision explanation — churning every identity uniformly would report “the rules have been updated” for rulebooks where nothing had. A rulebook that does carry corrections gets a new identity, which is the correct answer. The decision cache and the navigator cache use the same key, so a POST /rulebooks/{id}/fields edit now correctly invalidates both — the vocabulary is the one key component that is mutable in place, with no version cut, so before this a correction would have been masked by a warm entry for the whole TTL.
aethis-core
2026-08-02

Added

  • injection_phase gains post_submission (aethis-core#397). The lifecycle vocabulary ended at pre_submission, but some facts are produced by an assessor reviewing the evidence submitted with the application — they cannot exist before the submission that triggers them. Authoring such a field had no honest option: omitting the phase loses the due-date the axis exists to provide, and pre_submission asserts the fact is due before the thing that produces it. POST /rulebooks/{id}/fields rejected the authored value with a 422. Purely additive: the vocabulary stays closed (an unknown phase is still a ValidationError at authoring time and unknown_injection_phase at publish), every previously-valid value remains valid, and no stored content changes meaning. Consumers treat the phase as an opaque reported label — nothing compares, sorts or indexes it — so the new value needs no ordering decision.

Changed

  • INJECTION_PHASES and ELICITATION_OWNERS are removed. Each vocabulary existed twice — as a Literal (which rejects a bad value at authoring time) and as a hand-maintained tuple (which the publish gate renders into its error) — so adding a value to one and not the other produced a value that models accept and the gate rejects. The publish gate now derives each vocabulary from its type at the point of use, so the second name does not exist to fall out of step. No behaviour change. Deriving the tuples rather than deleting them was considered and rejected: an equality assertion catches the two copies diverging, but a contributor who reverts the derivation and re-types the values correctly ships green, leaving a hand-maintained copy for the next value addition to split. Internal constants only — both names were consumed solely by aethis_core.public.validation.field_elicitation, and neither appears in any request or response shape.
  • ElicitationViolation carries the vocabulary it applied, as a new allowed: Tuple[str, ...] field (empty when no closed vocabulary applies). Violation messages are byte-identical to before — verified across every violation case — and the routes that consume violations read field_key and message as they always did; this adds a structured reading of what the prose already said, so the guard for it need not parse prose back out.
aethis-core
2026-08-01

Added

  • Fact ownership and document recovery are now two independent, typed axes on every compiled field, and /decide reports what it withheld (aethis-core#389, workspace#835). 0.52.0 put elicitation_owner on the rulebook vocabulary; this puts both axes on FieldDefinition itself, so a ruleset carries them without a rulebook overlay, and makes the engine act on them rather than merely serve them.
    • elicitation_owner (applicant | caseworker | system | derived, default applicant) — WHO may supply the fact.
    • recoverable_from (list of evidence-capability names, default empty) — WHAT could supply it without asking. Independent of ownership: a date of birth is applicant-elicitable and passport-recoverable at the same time, so collapsing the two would either silence a good question or block a good document. Absence of a friendly question is explicitly not the safety mechanism.
    • injection_source / injection_phase — for a non-applicant fact, the named mechanism expected to supply it and the lifecycle point at which it is due.
    Both axes are carried through GET /rulesets/{id}/schema, GET /rulebooks/{id}/schema, next_question and every optimal_path entry, and are always emitted (including the applicant case) so a consumer reads askability instead of inferring it from an absent key.
  • pending_non_applicant_input on /decide — typed {field_id, owner, source, phase} for every required fact the applicant is not the source of which has not been injected. This is the half that stops the omission becoming a lie: with no applicant question left, an undetermined decision and a null next_question is otherwise indistinguishable on the wire from a finished case. The engine now selects only applicant-owned fields, and omits the rest from next_question, optimal_path and missing_fields — but they still constrain the solve, so the outcome stays undetermined until the value actually arrives.
  • Publish/promote gate for the elicitation contract (422 invalid_field_elicitation). Fails an unknown owner or injection phase, an applicant-owned field with neither a question nor a description (the runtime would otherwise fall back to a generated phrase and ultimately to the raw field id), and a non-applicant field that does not say where its value comes from or when. Unconditional, unlike the immigration-scoped grounded-why gate: a question the applicant cannot answer is a defect in any domain. Compatibility governs silence about ownership only. A field that declares no owner below schema v2 is legacy content, read as applicant-owned; from v2 it must say. Every other clause applies at every version — compatibility is a licence to omit the owner, never a licence to ship a field the runtime will have to invent a question for.

Fixed

  • GET /rulebooks/{id}/schema no longer discards a member field’s own authored question (aethis-core#390). The overlay fell straight through to description for any field with no vocabulary entry, so a field whose applicant-facing wording had been written during authoring served its internal engineering prose instead — 0.52.0 corrected only the one field named in the vocabulary. Precedence is now vocabulary → compiled question → description, matching what GET /rulesets/{id}/schema already did.
aethis-core
2026-08-01

Added

  • The Rulebook field vocabulary is now load-bearing: elicitation_owner + default_policy on RulebookFieldSpec, overlaid onto GET /rulebooks/{id}/schema (tda-server#901). rb.fields was written by set-fields, read back by the lock handlers, and consumed by nothing — it had no effect on the field universe served to a conversational agent. Nothing in compiled_fields said who may supply a fact, so every field looked equally askable, and a fact the applicant cannot know (the date the application is intended to be made — chosen by the supervising solicitor, not observed) was asked of them, invented, and then fed a statutory age exemption. elicitation_owner (applicant | caseworker | system | derived, defaulting to applicant) declares who may supply a fact. default_policy (currently only case_open_date) names a policy the runtime materialises into an ordinary stored fact at a known moment — deliberately not evaluated at decision time, which would make the same case decide differently tomorrow, break decision replay (decision_id / content_digest) and violate DS-41. The overlay is a left join: a vocabulary entry enriches a compiled field and never invents one; a compiled field with no entry passes through unchanged, so a partial vocabulary is safe. A rulebook with fields: [] — the deployed shape — is a behavioural no-op. Both keys are always emitted, so a consumer gets an explicit applicant rather than having to infer askability from absence. Additive on the wire: consumers predating the keys ignore them.
  • include_optimal_path on POST /decide, defaulting to true (aethis-core#374, step 1 of 3). optimal_path was computed unconditionally on every undetermined decision — a constraint-solver Optimize pass measured at ~4.4s of an ~11.4s evaluation on the real aethis/form-an rulebook, roughly 38% of a conversational turn, on a fallback that consumers read only when next_question is absent (so on the common path it is computed, serialised and discarded). It is now behind a request flag mirroring include_explanation / include_trace / include_graph_overlay, and joins the decision cache key exactly as include_composed_explanation does, so a flag-off entry can never be served to a flag-on caller. The default is true: this release is a behavioural no-op. Consumers opt out (step 2) and the default flips (step 3) as separate, later changes.

Performance

  • Version-pinned navigator cache entries no longer expire every 5 minutes (aethis-core#374). NAVIGATOR_CACHE_TTL applied a 5-minute expiry to an artifact the cache’s own contract calls immutable: a pinned rulebook composition folds every member ruleset_id and the outcome-logic hash into its key, so a re-publish mints a new key rather than changing what an existing one means — the TTL could never prevent a stale read, only force a ~0.43s solver recompilation of unchanged content. Production measured 83 rebuilds against 106 cache hits over 12h (a 44% miss rate). Pinned entries now take NAVIGATOR_CACHE_TTL_PINNED (24h) via the same per-entry TTL mechanism the decision cache already uses, and remain bounded by NAVIGATOR_CACHE_MAX eviction. Non-version-stable keys are unchanged.
  • Composed-explanation build 2.4x faster; no behaviour change (aethis-core#371). _check_condition_status — the constraint-solver core shared by the layered explanation, the graph overlay, and evaluate_subexpr_status — accumulated str(constraint) for every constraint it asserted, while the only consumer was len(...) in two diagnostic messages. str() on a solver boolean expression runs the solver’s pretty-printer over the entire AST, so every rendered string was built and discarded. Now a counter. Profiled offline against the real published aethis/form-an member bundles: composed explanation build 10.71s -> 4.56s, with 6.04s of the original attributable to the printer.
Decision provenance by reference, engine half — phase 2 (epic aethis-workspace#736 P2, aethis-core#360): the explanation is complete — composed (rulebook) explanations behind an opt-in flag, human-labelled supporting facts, and first-class “why ruled out”.

Added (P2)

  • include_composed_explanation request flag on /decide (#360). Opt-in composed (rulebook) explanation: when true on a rulebook_id request, explanation carries the layered per-criterion explanation across every member ruleset, with groups scoped <section>.<group> and per-criterion title / status / labelled supporting_facts / source_refs (plus per-member publish-resolved source_references, digest-bound to the exact member cut composed). Default false is load-bearing (spec Decision 7): downstream consumers auto-prefer a non-None explanation, so a default-on rulebook explanation would silently rewrite their user-visible narration. Composed mode omits decision_path — an all-sections composition has no single winning path, and naming one would be fabricated reasoning. The flag participates in the decision-cache key, so a warm entry written under one flag value can never serve the other. No-op on ruleset_id requests.
  • Labelled supporting facts (#360). Explanation supporting facts are now {field, label, value, display_value} — label joins the ruleset’s compiled_fields[].label (“Age at application”), and display_value is a deterministic human rendering (booleans as Yes/No, DATE ordinals as ISO dates, enum constants de-underscored) so renderers never need raw field ids. Additive: field and value are unchanged.
  • failure_reasons on not_satisfied criteria (#360). “Why ruled out” is first-class: each not-satisfied criterion carries the solver’s unsat-core reasons — the raw form (expression, plus a new fields list naming the input fields a condition references) AND a deterministic human string built with the same label join, free of raw field ids and raw solver syntax. Same-deploy determinism only; cross-version unsat-core stability is explicitly not claimed (the solver returns a core, not a canonical one).

Fixed (P2)

  • Failure-reason field/value mis-parse (#360). Unsat-core answer assertions were parsed back with rsplit("_", 1), fabricating non-existent field names for any underscore-bearing value (field eng.selt_provider answered ielts_ukvi was reported as field eng.selt_provider_ielts, value ukvi). Assertion meanings are now recorded at track time and looked up, never string-parsed; value now carries the actual typed answer rather than its string form.
Decision provenance by reference, engine half — phase 1 (epic aethis-workspace#736 P1, aethis-core#359): a decision can be tied to the exact rule content that produced it.Retained-bytes provenance, engine half (epic aethis-workspace#704 P1, aethis-core#350): every citation resolves at publish to retained bytes.

Added

  • content_identity on the /decide envelope (#359). One content-identity anchor for both request kinds: the leaf ruleset’s content_digest for a ruleset decide, or the version-pinned composition key for a rulebook decide (previously computed internally for the decision-cache key and never surfaced, so a rulebook decision had no exposed content identity at all). Replay-critical: a caller that retains it can later verify an explanation was recomputed against the same rule artefact, and refuse it if that artefact has since been superseded. The two kinds bind at different strengths. A leaf content_digest is a true content identity (a hash of the compiled bundle). A rulebook composition key is a composition identity: it binds version-pinned member ids plus the outcome logic, not the members’ compiled bytes. Under the converged promotion model each cut pins a new bundle id, so the two coincide in practice — but a legacy path mutating content under a stable bundle id would leave the value unchanged across that change. Content-level binding for compositions is #39; until it lands, a rulebook content_identity is composition identity. Null is meaningful. A rulebook whose members are not all version-pinned has no stable composition identity, so the field is null rather than a fabricated value — such a decision is not verifiably replayable, and a consumer verifying provenance must refuse it instead of reading null == null as a match. The value rides the cached payload like every other artefact-identity field; because the decision-cache key embeds content_identity, a warm hit can only be an entry written under the same identity, so a cache hit reports the same anchor as the cold call that populated it.
  • section_results[].ruleset_id (#359). Rulebook per-section diagnostics now name the member ruleset that produced each section, resolved from the same composition the navigator compiles. The field was declared but never populated. A section prefix with no matching member stays null — an honest gap, never a guessed id.
  • Artefact citation targets (#350). source_targets entries now take exactly one of url or artefact_source_id (both/neither → 422). An artefact target resolves against the caller’s uploaded project source (ProjectSource + ProjectSourceBlob, scoped by tenant + project) with zero network calls: row/blob/re-hash digest and size agreement is verified, the verbatim quote is checked against the retained bytes (PDF page-aware), and the emitted reference is schema v2 — target_kind: "artefact", artefact_project_id, artefact_source_id, url = the authenticated /raw download route. Distinct fail-closed reason codes: artefact_not_found, artefact_cross_scope, legacy_source_no_artefact, artefact_integrity_mismatch, artefact_unsupported_type (plus the existing quote_not_found / digest_mismatch).
  • Snapshot-on-fetch for URL citations (#350). The exact bytes fetched for a URL target are retained in the new content-addressed citation_snapshots collection (deduplicated by sha256: digest, 10 MB fetch cap unchanged); newly emitted v1 references carry the additive optional snapshot marker (the snapshot digest). Previously stored v1 references are served byte-for-byte unchanged — absent additive fields are omitted, never emitted as null.
  • Artefact public-exposure guard (#350). Reason code artefact_reference_public_visibility_forbidden on every lane to public: public-bound publish with artefact targets, ruleset visibility flip when any stored cut carries a v2 reference, rulebook visibility flip when any member does, and member attach / promote-to-live into an already-public rulebook. Private uploaded sources can never become anonymously resolvable.
  • Released-consumer parse-compat fixture (#350). tests/contract/fixtures/explain_v2_artefact_response.json is generated through the engine’s own response models, and CI proves the released PyPI aethis-cli and aethis-sdk parse/render a v2-bearing response without failure.
aethis-core
2026-07-28

Fixed

  • Every /generate call failed on 0.49.1 (#356). GenerationInputs grew a sixth field (source_document_ids, part of the #331 provenance wave) while _run_generation_inline still unpacked five positionally, so authoring generation raised ValueError: too many values to unpack (expected 5) immediately on staging and production. Now consumed by attribute; the same function already read inputs.source_document_ids elsewhere, so the unpack was the single stale site. Caught by a live canary, not by CI — a positional unpack of a NamedTuple is type-checkable but was not covered by a test that ran the inline generation path.
aethis-core
2026-07-28
Released as 0.49.1. 0.49.0 was bumped in the repo but never tagged — the deploy script bumped again at tag time, so no v0.49.0 exists. These notes describe the 23 commits in v0.46.3..v0.49.1.
Public developer release (epic aethis-workspace#643) — production safety boundary and truthful contracts. This is the version the public docs at docs.aethis.ai describe. Until it is live, api.aethis.ai serves 0.46.3 and the differences are listed on Deployed contract status.Minor rather than patch because three documented behaviours change for existing callers — see Changed below. Pre-1.0, a minor bump carries breaking changes.

Added

  • Immutable published ruleset versions (#335). PublishedRulesetVersion documents are write-once (unique idempotency_token, unique (ruleset_id, content_digest) backstop), fronted by a CAS-protected publication head with a generation counter and an append-only PublicationEvent stream (publish / head-move / retire / refs-enriched / corrupt-record / quarantine). Republishing mints a new version and moves the head; it never mutates a prior version, so a version id remains replayable forever. One canonical version primitive — both the legacy publish path and promote_ruleset_to_live go through cut_published_version.
  • Real resolved identity on every leaf response (#335). /decide, /schema and /explain now return the resolved immutable ruleset_id (never the caller’s slug), the published ruleset_version, and a content_digest (sha256:<hex>). The digest is folded into the navigator and decision cache keys, so a republish under a stable slug cannot serve prior content from any cache tier — including from a Cloud Run replica that never saw the publisher’s invalidation.
  • Publish-validated source_references[] on criteria (#335). A public citation is a self-locating reference validated at publish time (content digest plus verbatim-quote occurrence), not an opaque string.
  • Documented response truth table (#335). aethis_core/public/contracts/decision_truth_table.py is the canonical machine-readable table, mirrored in docs/public-decision-truth-table.md and drift-guarded by tests.
  • Route- and method-scoped anonymous CORS (#333). RouteScopedCORSMiddleware grants any origin (no credentials) on the documented anonymous evaluate/read routes only; everything else is restricted to trusted first-party origins, with localhost off in production.
  • X-Aethis-Records-Omitted on catalogue reads (#335), so a truncated listing announces itself instead of looking complete.
  • Content-addressed source artefact retention (#331), with cited-source deletion guarded so a citation can never dangle.

Changed

Three changes are visible to existing callers:
  • A positive verdict can no longer coexist with blocking input errors (#335). Non-empty field_errors forces decision: "undetermined" as the last word before the response is built, and clears misleading guidance. Per-criterion facts (explanation.groups[].status, graph_overlay.nodes[].overlay.status) still report what the engine established from the valid subset of inputs, so a criterion may read satisfied inside a forced-undetermined response. decision is the only outcome field; never re-derive an outcome by aggregating group statuses.
  • An undefined top-level request key is now rejected (#335). DecideRequest is extra="forbid", so a misspelled or unsupported key returns a structured 422 extra_forbidden instead of being silently ignored with a 200.
  • Empty-extraction source uploads fail loud (#327). An oversized (>500-page), image-only or undecodable upload returns 422 and stores nothing, instead of storing an empty source with extraction_error set and returning 201 — which let generation run ungrounded against a source the author believed was loaded. A partial extraction that still yields text is unaffected: stored, with extraction_error surfaced.

Security

  • Closed environment predicate (#333). is_production / is_deployed normalise prod and production; an unsupported value refuses to boot, and DISABLE_AUTH=true refuses to construct on any deployed environment. The deploy pipeline sets _ENVIRONMENT=prod, which the previous environment == "production" comparison never matched — so the production service was silently running without the DISABLE_AUTH guard or HSTS, and admitting localhost CORS.
  • Fail-closed auth (#333). The all-scope mock principal is minted only when auth is disabled and the environment is non-deployed.
  • Anonymous throttling keyed off a trusted proxy hop (#333). Client identity comes from the rightmost trusted-appended X-Forwarded-For entry, so a forged header rotation can no longer evade a 429. Adds per-IP and global anonymous budgets, an anonymous concurrency cap, costed expensive options, and a pre-parse body cap (413 before JSON parsing).
  • Anonymous decision logging is metadata-only (#333). Hashes and aggregate metadata only — never raw field_values or caller_ref — through a single choke point with a seeded-canary test.
  • internal_only honoured on the authenticated cross-tenant path (#333). visible_to now requires visibility=public AND internal_only != true, so an internal_only ruleset wrongly flipped public is not readable by any API key on a public engine. A loud boot-time invariant scan backs it.
  • No real tenant identifier in the public OpenAPI document (#337). The caller_ref example named a real pilot firm; it is now a neutral placeholder, guarded by a test that asserts on the serialized document.
  • A rulebook can no longer compose another tenant’s private ruleset (#334). ruleset_refs were stored verbatim from the request body, and every path that later resolved a member — the navigator (so the decision itself), the /decide graph overlay, and the rulebook schema / fields / graph / activate reads — looked it up with no tenant or visibility predicate. A tenant could therefore point its own rulebook at another tenant’s private compiled-ruleset id (or collide on a non-namespaced section_id), call /decide on its own rulebook, and receive a decision computed against the other tenant’s rules plus their criteria text. Closed at two independent gates: refs are validated against visible_to on create and update (422 ruleset_not_visible), and all six resolution sites are scoped to the rulebook owner’s visibility — including a re-assertion on the compiled-bundle cache-hit path, since that cache is keyed by the compiled-ruleset id with no tenant component and would otherwise launder an unauthorised read. Composing a public aethis/* ruleset into a private rulebook is unaffected and covered by tests.
    Tier-2 tenancy — this must carry an independent adversarial review before the tag is cut. It was authored and self-tested in one session, which adversarial-review-discipline does not accept as sufficient for a tenancy change. Do not treat its presence in these notes as review having happened.

Changed (continued)

  • Anonymous rate limiting moved to layered rolling windows (#343). The anonymous lane had been left on the pre-#552 shape — one counter per calendar day, and Retry-After: 86400 on every breach — because epic #552 re-sliced the authenticated path and explicitly declined to re-tune anonymous. It now meters three rolling windows per class: burst (the current minute), sustained (rolling 1h) and daily (rolling 24h), read from one minute-bucket granularity in a single aggregation. The tightest breach is the one reported and Retry-After names the shortest window that actually clears, so a breach costs at most one window rather than lasting until UTC midnight. read (schema/list, near-free) now out-budgets decide (the compute), where both were previously capped at 500/day. The shared global cap gets the same treatment — it could previously 429 every anonymous caller until midnight. Numbers are owner-tunable policy data; this is a loosening of an abuse control and may warrant a PILOT_RELAXATIONS entry.

Fixed

  • A rulebook member’s source references now survive publish → promote (#336). Publish-time source resolution was gated on rulebook is None, so a ruleset published as a rulebook member never resolved its declared citation keys through the publish path at all, and the cut that records them sat in the same branch. Members now resolve like any other leaf and cut a version carrying the references, deduping with the promote-time cut via the existing idempotency contract. The D2 traceability claim for aethis/uk-fsm/child-eligibility remains unverified until a live publish → promote → /explain run confirms it end to end.
  • Upload success-path test coverage restored (#339) — the invariant that the empty-extraction guard does not fire on a partial decode had no working coverage, because the only test exercising it never ran.
  • generate-and-test timeout raised to a tunable 270s below the Cloud Run ceiling (#326).
  • DSL_TEMPLATE operator list derived from the Operator enum, with drift-guard tests (#328).

Deploy notes

  • Blocking, before the tag: #334 needs an independent adversarial review. Resolved — the review ran and found real holes. An independent fresh-context reviewer returned BLOCKING on two counts: (a) the fix had no regression coverage at its highest-value seam — deleting one line (tenant_id=rulebook.tenant_id) reopened the vulnerability on /decide byte-for-byte with the entire suite green, as did deleting the cache-hit probe call and de-scoping the probe’s query; (b) the “all six member-resolution sites are scoped” claim was false — a seventh site (_resolve_section_names) leaked another tenant’s private ruleset name to the anonymous rulebook catalogue. Both closed in #348, which also fixed a 24h cache-revocation window, a falsy-vs-is None inconsistency across five sites, and two latent traps. The three surviving mutations are now each caught by a named test. Shipped in v0.49.1. Independence obtained: different-session + adversarial framing; not different-provider, which adversarial-review-discipline asks for on Tier-2 — a second review from another model family would still be additive.
  • Unverified, before the claim: #336 restores the mechanism but the D2 source-traceability claim on aethis/uk-fsm/child-eligibility needs a live internal-key publish → promote → /explain run. Until it passes, that showcase is excluded from the traceability claim or the claim is narrowed.
  • Owner-unreviewed policy: #343’s anonymous limits were chosen without a sign-off on the numbers. They are policy data; re-tune in one line if the exposure is wrong for launch.
  • Every acceptance item requiring a live revision — check-public-deploy-security.py against production with --burst / --concurrency-probe / canary, plus the TTL-index, backup-retention and raw-canary-absence attestations — is deferred to the approved production deploy and consumed by epic #643 P10. The script emits these as pending-attestation items with the exact canary token to search for.
  • After the tag is live, refresh the canonical demo bundles (aethis-examples: make rebuild-canonical-bundles then make snapshot-canonical-bundles) — this release changes decision-envelope shape.
  • Update mintlify-docs/api-reference/openapi.json info.version and delete reference/deployed-contract.mdx (plus the caveat snippet that links to it) once engine_version reports 0.49.0.
aethis-cli
2026-07-27
Publish with citations — and be able to cite a document you hold, not just one the internet happens to host.
  • feat: aethis publish --source-targets <file> resolves a ruleset’s citation keys. A YAML or JSON targets file maps each citation key your criteria declare (source_refs) to the document it cites: title, authority, licence, and the verbatim quoted text. The engine verifies every quote against the source bytes at publish time and rejects the publish if any citation fails — there is no half-cited ruleset.
  • feat: a citation can point at a file you uploaded, not only a public URL. An entry naming file: is uploaded to the project and cited by its source_id, so the rules can cite the very documents they were generated from. The engine resolves it from retained bytes with no network call at all.
  • feat: an identical file is never uploaded twice. File targets are matched by sha256 against the project’s existing sources and reused when the bytes are already there — including two entries in the same run naming byte-identical files. The API does not deduplicate uploads, so re-running a publish previously grew the project a duplicate source per citation.
  • feat: a malformed targets file costs no round trip. Exactly-one-of url/file, a readable file, an https:// URL, the required title/authority/licence/quote.exact, and unknown fields are all checked locally — every problem in the file reported at once, before the first API call, so nothing is uploaded against a targets file that was never going to publish.
  • feat: an uploaded-artefact citation is never rendered as though it were a public link. aethis decide --explain and aethis explain label these references as an uploaded snapshot verified at publish, state that the download is authenticated and needs a key with projects:read on the project, and resolve the engine-relative download path against the host you called — while a URL citation keeps reading as the public link it is. aethis publish reports the same distinction per target as it resolves them.
  • feat: aethis can list a project’s uploaded sources (AethisClient.list_sources), which is what makes the digest comparison above possible.
  • fix: a rejected citation now says which one and why. Publish-time citation resolution is fail-closed and the API itemises every failing key ({source_id, reason_code, message}), but the CLI collapsed the whole envelope to its summary line — an author with three citations and one wrong quote learned neither which key failed nor what was wrong with it. Every itemised failure is now printed under the error. Applies to any endpoint returning a failures list, not just publish.
  • feat: citation targets that never landed are reported, not silently dropped. The engine only resolves the citation keys the compiled ruleset actually declares, and ignores the rest — so a mistyped key published “successfully” with zero citations attached, after uploading the files. After a publish with --source-targets, the CLI reads the published ruleset back and says how many targets landed, naming any that did not. A failed read-back degrades to an honest note; it never turns a successful publish into a failure.
  • feat: a failed publish says your uploads are still there. Publishing is fail-closed but the uploads that preceded it are not rolled back, and re-running reuses them by digest rather than duplicating them. The error now says so, instead of leaving the state of the project a guess.
  • fix: a duplicate citation key is rejected instead of silently overwriting. YAML and JSON both let the last definition of a repeated key win; in a citation manifest that quietly discards a document and publishes the other one into an immutable ruleset. Both formats now refuse duplicate keys (at any depth in the file).
  • ci(publish): the downstream-unstick sweep now reaches every consumer repo. The unstick-downstream job searched a single owner, so a PR carrying an aethis-needs: aethis-cli marker in a repo under a different owner was never found and sat as a draft indefinitely after the release it was waiting for went out. It now queries each consumer repo individually with gh pr list --repo, which is genuinely repo-scoped, and matches the marker against the PR body returned by that same call (one request per repo instead of a search plus a fetch per hit). A repo the token cannot read is reported as a warning instead of failing the whole sweep. No package or runtime change.
aethis-sdk-python
2026-07-26
Makes the SDK a safe, immutable release component for the public developer release (epic aethis-workspace#643, P9 / aethis-sdk-python#29). Three classes of “looks fine, isn’t” are closed at the type level, and the release itself now carries verifiable integrity evidence.

Replay identity: absence no longer looks like a version

  • breaking (behavioural): DecideResponse.ruleset_version is str | None and no longer defaults to "unknown". The engine reports an unresolved version as the literal string "unknown" (a rulebook call, or an artefact published before immutable versions); the SDK also defaulted to that string, so a caller writing an audit record got a plausible-looking version whether or not anything had been resolved. Every unresolved sentinel ("unknown", "", "none", "null", "n/a", case-insensitive) now normalises to None on ruleset_version, content_digest, ruleset_id, engine_version, decision_id and inputs_hash. Code reading response.ruleset_version gets None where it previously got "unknown".
  • feat: content_digest on DecideResponse, and ruleset_version + content_digest on SchemaResponse. The resolved immutable identity aethis-core stamps on /decide, /schema and /explain (aethis-core#330).
  • feat: ReplayIdentity / ContentIdentity + require_replay_identity() / require_content_identity(). These return a complete identity or raise AethisReplayIdentityError naming exactly which parts are unresolved — so recording an incomplete audit reference is an explicit act, not a default. The soft accessors replay_identity / content_identity return None instead of raising.

Blocking errors cannot become a completed or positive result

  • feat: DecideResponse.blocking_errors (always a mapping), .has_blocking_errors, .is_terminal, .raise_for_blocking_errors(). The engine suppresses next_question while blocking field_errors are outstanding, so a blocked response is byte-shaped like a finished one on that field. is_terminal is the honest check.
  • feat: the parse boundary refuses a self-contradicting envelope. A 2xx reporting eligible/not_eligible beside non-empty field_errors — or an embedded copy in explanation.decision / trace.status that contradicts the headline — raises the new AethisContractViolation rather than becoming an object a caller acts on. Enforced in the model, so it holds on the sync client, the async client and the sessions alike.
  • feat: SessionStatus gains field_errors, replay_identity, .blocked, .is_complete and .raise_if_blocked(), with a constructor invariant that makes a positive-and-blocked status unconstructible. Sessions gain blocking_errors() and is_complete() (sync and async).
  • feat: new AethisFieldErrors exception carrying .field_errors, raised by the opt-in raise_for_blocking_errors() / raise_if_blocked() guards.

Typed source provenance

  • feat: SourceReference + SourceQuote models — the publish-validated citation contract (source_id, title, authority, HTTPS url, locator, source_version, source_date, content_digest, licence, verified_at, verbatim quote, self-locating deep_link, schema_version), returned identically by both explanation surfaces. Unknown fields are preserved so the additive schema_version evolution cannot break a pinned consumer.
  • feat: get_explanation(ruleset_id) (sync + async) returning the typed ExplainResponse, with resolved identity and typed references per criterion. explain() still returns the raw dict for existing callers.
  • feat: DecideResponse.decision_explanation + .source_references() parse the /decide explanation into DecisionExplanation. Note the two surfaces differ: /explain returns a flat criteria array, /decide nests criteria under explanation.groups[].criteria[]. They share the DTO, not the envelope — the SDK models them separately and the tests assert the distinction.

The two access boundaries are labelled

  • feat: AethisError.boundary is "evaluation" or "authoring" on a 401/403, and the exception message now names which door was closed — no-key evaluation (/decide, /rulesets, /schema, /explain) versus invite-only authoring — plus the access-request URL. README and examples carry the same labelling before either path.

Release integrity and hermetic install evidence

  • feat: scripts/release_integrity.py emits the tuple a release candidate is pinned on — (package, version) → exact sdist/wheel sha256 → source commit/branch/clean-state — and re-verifies it against local files or against what PyPI actually serves. Wired into publish.yml before publication (--require-clean) and after (--verify-registry).
  • feat: scripts/hermetic_install_check.py installs the exact artefact into a throwaway world — temporary HOME/XDG_*/cache, every AETHIS_* and provider key unset, empty cache on first install, no alternate index — then runs an offline smoke that parses captured engine payloads through the installed package, and a poisoned-artefact negative control that must fail. New hermetic CI job across ubuntu/macos × Python 3.11/3.12/3.13.
  • feat: scripts/capture_engine_fixtures.py records the fixtures under tests/fixtures/ from a live engine (anonymously, against a public showcase ruleset), including the engine’s own JSON Schemas, so the mocked suite is tested against real wire payloads rather than hand-written approximations.
  • chore: Python 3.13 added to the classifiers and the CI matrix. jsonschema added to the dev extra (test-only; the shipped package is still just httpx + pydantic).

Review-wave fixes (same release)

  • fix: get_source() reported the wrong access boundary. /rulesets/{id}/source sits under the /public/rulesets prefix but is key-required behind a scope external keys are not issued, so the prefix match labelled a 401 there "evaluation" and told the reader to go looking for a ruleset-visibility problem for a door that will never open. Key-required sub-paths are now excluded from the evaluation prefix, and both get_source docstrings say so. Prefer get_explanation(), which is anonymous on a public ruleset and returns the same SourceReference DTO.
  • fix: content_digest is validated against ^sha256:[0-9a-f]{64}$. md5:…, a truncated sha256:beef, non-hex, uppercase and bare-hex values now normalise to None rather than being carried into an audit record — the same “looks like a value, isn’t” class as the "unknown" version.
  • fix: --require-clean passed vacuously when git provenance was unreadable. source_provenance() returns dirty: None (unknown) when git cannot be read, and the gate tested it for falsiness — so with no .git the script exited 0 while recording commit: null. Provenance problems are now checked positively (dirty is not False, plus a missing commit in its own right).
  • fix: the poisoned-artefact control was vacuous on every CI runner. It flipped the final byte, which lands in the end-of-central-directory comment field; strict zip readers reject it, lenient ones scan backwards and install happily. It is replaced by two controls: a digest control (a valid, installable substituted wheel that the real verify_files must reject — deterministic, and the layer that actually protects users) and an installer control (a wheel whose compressed stream is corrupted, asserted positively on stage == "uv pip install" plus stderr that names the corruption, rather than on the absence of one unrelated string).
  • fix: publication is gated on --verify-files immediately before the upload, which is irreversible, in addition to the post-publish registry check.
  • fix: the README session loop no longer demonstrates the bug this release exists to prevent. It looped on next_question() is not None; it now loops on status(), with a table of the four states that loop conflated.
  • fix: ruff’s rule selection is pinned (ruff>=0.6.0,<0.17 plus an explicit [tool.ruff.lint] select). CI installed ruff unpinned, so 0.16.0’s wider default turned the lint gate red on 61 errors — before pytest ran at all. Mirrors aethis-cli#90. CI also now lints examples/, which it had never covered.
  • fix: captured OpenAPI examples are stripped, and a test guards the fixtures. The engine’s caller_ref example named a real pilot firm, which a verbatim capture committed to this public repo. examples carry no structural information and validation ignores them, so they are no longer captured; a standing test fails on any internal tenant name or immigration term in tests/fixtures/.
  • fix: the wheel is now byte-reproducible. Three builds of the same clean tree produced three different digest pairs, so the tuple’s “which commit produced these bytes” leg was an attestation nobody could re-derive. SOURCE_DATE_EPOCH is now set from the commit timestamp in both build workflows, a CI step rebuilds the wheel and fails if the digest moves, and the tuple records precisely what is verifiable — the sdist is still not reproducible (setuptools varies the archive), and says so rather than implying otherwise.
  • test: the two invariants that no ordinary test could reach are now covered. is_terminal’s blocking-error clause and SessionStatus.is_complete’s not blocked clause are each masked by the parse validator and the constructor invariant respectively — so deleting either left the suite green while removing the last guard on a bypass route. Both are now exercised through model_construct / object.__setattr__ / dataclasses.replace, and verified to fail when the clause is removed.
aethis-cli
2026-07-26
Safety and provenance for everything the CLI reads back from the API.Versioning note. This release changes two documented behaviours — aethis explain --output json emits the whole envelope rather than the bare criteria array, and a blocked evaluation now exits 3 where it previously exited 0. Under strict SemVer a breaking change is a major bump; taken as a minor here because the package is pre-1.0 (0.x), where the published rule is that minor carries breaking changes. Recording the call explicitly rather than leaving it to be inferred: both changes replace behaviour that was unsafe (a script could not tell a rejected input from a decision), which is why they ship rather than waiting for 1.0.
  • feat: a rejected input can never look like a result. When a decision response carries blocking field_errors, aethis decide and aethis rulebooks decide print the rejected inputs instead of a verdict and exit 3 (new exit code: 0 decided, 1 call failed, 3 inputs rejected). JSON output reports "decision": "undetermined" and records the block under aethis_cli_contract. The CLI enforces this rather than trusting it: if a server ever returns eligible beside blocking errors — a stale deployment, a caching proxy, a third-party API-compatible server — the contradiction is overridden, reported, and never rendered as success in human output, in JSON, or through the exit status. New aethis_cli.contract module owns the rule.
  • feat: immutable identity on every decision surface. Human output gains a Ruleset identity block (ruleset id, published version, sha256: content digest, engine version, decision id, inputs hash) so a decision can be reproduced or audited later. A ruleset_version of unknown — which a published ruleset must never report — is called out as unresolved instead of printed as though it were an identity.
  • feat: supporting sources are shown, and kept separate from the rules. Where a ruleset publishes validated source references, aethis decide --explain and aethis explain render them under their own Sources heading: document title, authority, locator, the verbatim quoted text, deep link, licence, verification time and source digest. A reference that arrived incomplete is flagged rather than rendered as a confident-looking citation. When a ruleset publishes none, the output says so rather than showing nothing.
  • feat: output distinguishes result, logic trace and source. The three now have their own headings, and the logic trace is labelled as explanatory — per-criterion statuses answer “what is true of this criterion”, never “what may I act on”.
  • change: aethis explain --output json now emits the whole envelope (ruleset_id, slug, ruleset_version, content_digest, criteria) instead of the bare criteria array. Provenance on the machine-readable path is the whole point of the identity contract; a script that consumed the old shape needs .criteria.
  • feat: undeclared /decide request options are refused locally. The API rejects an unknown top-level request member with a 422 rather than ignoring it; the CLI now names the offending option before spending a round-trip, and renders the server’s validation envelope readably if one is ever returned.
  • feat: the capability boundary is visible before you hit it. Root help, decide/explain help, the README and every auth-required error now state plainly that evaluation needs no account and no key, and that authoring is invite-only (with the access link).
  • feat: release integrity and hermetic first-install evidence. scripts/release-integrity.py binds the exact published bytes to the commit they were built from (version + sdist/wheel sha256 + source commit) and can re-check that against the files the registry serves. scripts/hermetic-install-check.py proves the CLI works for someone who has never run it: temporary HOME/XDG/config/cache, no Aethis or provider credentials, empty-cache install from one source only, across supported runtime/OS/architecture — with a poisoned-cache negative control that must fail. Both run in CI on every PR (Linux + macOS, Python 3.11/3.12/3.13) and on every release, before and after publication.
  • feat: the guard matches the engine’s own forcing sweep exactly. When a response is blocked, all five embedded copies of a terminal verdict are scrubbed — top-level decision, explanation.decision, explanation.decision_path, trace.status, trace.path — so a blocked result can never be printed above a green “Satisfied by: …”. A --json <fields> projection can no longer drop the aethis_cli_contract record either.
  • fix: a non-JSON response body is an error, not a traceback. A 2xx carrying HTML (an intermediary’s error page, a truncated body) now surfaces as one readable API error instead of a JSONDecodeError escaping mid-command.
  • fix: documented commands that would not run. --output is a root option and must precede the subcommand; five documented invocations had it after (including the README’s flagship shell-gate example, whose else branch therefore misreported). All fixed, and a test now resolves every documented invocation against the real command tree so this class cannot come back.
  • test: the contract oracle is itself gated. scripts/mutation-check.py breaks the contract sixteen different ways — deleting each scrub site, redefining the blocking exit code, making the blocking predicate always-false, unguarding the rulebook surface — and requires the suite to go red for every one. It runs in CI.
  • test: contract fixtures are captured, never hand-written. The new tests run against payloads recorded from a live engine (terminal decisions, each class of blocking input error, an incomplete evaluation, the 422 for an undeclared request member, the explain envelope) plus source-reference DTOs serialised by the engine’s own model. Regenerate with scripts/gen-contract-fixtures.py; provenance is documented in tests/fixtures/contract/README.md.
  • fix: no spurious traceback when a login callback is cancelled or times out. The OAuth callback server closed its socket while its background thread was still waiting on it, so the thread died on ValueError: Invalid file descriptor: -1 and printed an unhandled-exception traceback after aethis login timed out. The serving thread is now signalled before the socket closes, and treats a closed socket as its exit condition rather than an error.
aethis-mcp
2026-07-25
  • security: every tool’s server-supplied free text is fenced as untrusted data. JSON-passthrough tools (list/discover/schema/graph/rulebook and friends) now wrap their whole response in a single <api_response> fence with the untrusted preface, so a server- or tenant-authored free-text field (name, description, domain, message, …) can no longer smuggle instructions to the model via the JSON blob. Prose tools keep their per-field fences. A new deterministic serializer-coverage test drives every tool with taint sentinels and fails if any free-text leaf is ever emitted outside a fence. Closes the untrusted-JSON-passthrough gap (aethis-mcp#45).
  • security: capability annotations + containment. Every tool now carries MCP annotations (readOnlyHint / destructiveHint / openWorldHint) derived from a single capability registry, so a host renders correct read/destructive hints and can gate approval on the mutating tools. A test enforces that no no-API-key (anonymous) tool has any mutation capability, and that the registry matches which handlers actually require a key.
  • release: server.json is derived from package.json and drift-guarded. server.json (the official MCP Registry record) now tracks the package version source of truth; npm run check:server-json (run in the test suite/CI) fails on drift. Corrected the stale server.json version (0.5.1 → current).
  • release: generated tool inventory. tool-inventory.json is a generated, drift-guarded listing of the tool surface, used to verify a fresh install and as the source of truth for the published tools reference.
  • release: gated, evidence-producing publish pipeline. An unprivileged build stage produces one immutable tarball with a sha256 digest, a CycloneDX SBOM and a candidate manifest; npm and the official MCP Registry are then published via separate protected environments (named reviewer), workflow-bound OIDC and no stored token. Post-publish verification and a clean-environment fresh install must both confirm the exact name/version before a release reports success. See docs/RELEASE.md.
aethis-mcp
2026-07-25
  • docs: align Simpson paper citations with v3.13 (issue #53). The construction-insurance demo now cites the current paper version (v3.13, 2026) instead of v3.11; removes any presentation of the withdrawn GPT-5.4 low-reasoning-effort 7/11 figure as a live result (the v3.8 withdrawal note remains as historical context); and removes configuration-level API detail (parameter names) from the benchmark methodology text. All real benchmark numbers are unchanged.
aethis-sdk-python
2026-07-21
  • feat: usage() + client.rate_limit — rate-limit budget + headers. New Aethis.usage() / AsyncAethis.usage() return a UsageResponse (per-operation-class used/limit/remaining/reset over the rolling 24h window + a 7/30-day rolling summary) from GET /api/v1/public/usage. Every response’s X-RateLimit-* headers are now parsed onto client.rate_limit (a RateLimit model: operation_class/limit/remaining/reset), so a consuming app can read its remaining budget — especially generate (the scarce LLM class) — without a separate call. New models UsageResponse, ClassUsage, RollingUsage, RateLimit, all exported. (epic aethis-workspace#552)
  • Requires aethis-core’s /usage + X-RateLimit-* surface (epic #552 P2); the public release of this version is held until that is live on api.aethis.ai.
  • ci: cut a GitHub Release on publish. The publish workflow now creates a GitHub Release for each just-published tag, using that version’s CHANGELOG.md section as the release notes, so the “watch → releases” subscribe channel stays current automatically. Idempotent (create-or-skip on an existing Release) and --verify-tag (never mints a synthetic tag). No package/runtime change. (epic aethis-workspace#526)
aethis-mcp
2026-07-21
  • feat: aethis_usage tool. Reports the caller’s rate-limit budget per operation class (decide / generate / author / read / keys / admin) over the rolling 24h window — used, limit, remaining, reset — from GET /api/v1/public/usage, so an agent authoring inside Claude Code / Cursor / Windsurf can see and report the developer’s remaining generate budget before a 429. New AethisClient.usage(); tenant-scoped (requires an API key). (epic aethis-workspace#552)
  • Requires aethis-core with the /api/v1/public/usage endpoint live (epic #552 P2). The public npm release of this version is held until that endpoint is live on api.aethis.ai.
aethis-cli
2026-07-21
  • feat: aethis usage — show your rate-limit budget per operation class (decide / generate / author / read / keys / admin) as a table: used / limit / remaining / reset, over the rolling 24h window. generate (LLM rule generation) is the scarce class; browsing and status polling (read) are effectively unlimited. --json/piped emits the raw /usage payload. New AethisClient.usage().
  • feat: remaining generate budget after aethis generate. The CLI now reads the X-RateLimit-* response headers (captured on every request as AethisClient.last_rate_limit) and, after a generation is queued, prints “N generations left in the current 24h window” — so a 429 is never the first signal. The line turns yellow at ≤5 remaining.
  • Requires aethis-core with the GET /api/v1/public/usage endpoint + X-RateLimit-* headers live (epic aethis-workspace#552, P2). The public release of this version is held until that surface is live on api.aethis.ai.
aethis-core
2026-07-21
Metering & rate-limit revamp (epic aethis-workspace#552) — P4 over-limit rate.

Added

  • Per-key over-limit rate is now persisted and queryable (aethis-core#320 follow-on). A new over_limit_count field on the hourly rate_limits bucket is incremented whenever a request is observed over its class limit — in report-only mode too, since the observation is the oracle the Bridge tunes limits against (not just a log line). The write is best-effort and fully exception-isolated: a failure can never alter the rate-limit decision or 500 the request. GET /api/v1/admin/usage/overview now returns over_limit_24h per (key, class) — the rolling-24h over-limit count — so the console usage dashboard (ok_swift#665) and the approaching-cap alert (godseye#46) can show the 429-rate metric alongside proximity-to-cap.
aethis-core
2026-07-21
Metering & rate-limit revamp (epic aethis-workspace#552) — P4 admin primitive.

Added

  • GET /api/v1/admin/usage/overview (aethis-core#320) — internal-gated (admin:read scope AND internal==true, same stack as the rest of /api/v1/admin/*), read-only, cross-tenant. Returns, per (key_id, tenant_id, tier, class), the rolling-24h used joined against the tier×class limit with remaining and pct_of_limit, ranked by proximity to cap descending (by pct_of_limit, never raw used — so a near-cap low-tier key is never lost behind a high-volume internal one) so “approaching cap” is the top of the list. Filters: tier, key_id, class, min_pct (the alert threshold), limit; total_matching reports the pre-limit count so a caller knows when the display was clamped. The rolling-24h used is the same hourly-bucket sum the enforcement path and GET /public/usage use, so the overview can never disagree with per-key usage. This is the shared primitive both ok_swift#665 (console usage dashboard) and godseye#46 (approaching-cap alert) consume. Stays in the live /openapi.json but out of the published public contract. Over-limit history (429-rate) is a deliberate follow-on: the rate_limit_over_limit event is log-only, so a history feed needs a metering write-path change.
aethis-core
2026-07-21
Metering & rate-limit revamp (epic aethis-workspace#552) — reaches production.

Added

  • GET /api/v1/public/usage — calling-key-scoped usage per operation class (rolling 24h used/limit/remaining/reset, plus 7-day and 30-day rolling totals). Deliberately unmetered.
  • X-RateLimit-* response headers (Class/Limit/Remaining/Reset) on metered endpoints — forward visibility so callers can see their budget without waiting for a 429.

Changed

  • Rate-limit counters re-sliced into six operation classes — decide, generate, author, read, keys, admin (one counter is both the limit and the usage metric). Policy is expressed as tier×class data.
  • Meter the scarce thing: only generate (LLM-backed) carries a real ceiling. The old shared projects authoring bucket (2000/day — the squeeze that made a large rulebook “hit the 500 limit”) is split into author (20k/day) and read (100k/day), so ordinary authoring no longer competes with generation.
  • Rolling 24-hour window (hourly sub-buckets) replaces the calendar-day reset.

Notes

  • generate enforcement ships report-only (RATE_LIMIT_ENFORCE env, default off): the scarce-class limits are logged (rate_limit_over_limit event) but do not reject, pending owner tuning on real 429-rate evidence. Every other class is preserved-or-loosened vs 0.46.2, so no traffic that succeeds today is newly rejected.
aethis-mcp
2026-07-20
  • Startup update-check nudge. On startup, the server checks the npm registry’s latest version for aethis-mcp and, if a newer release exists, writes a one-line notice to its stderr log (visible in your MCP host’s server logs) pointing at the Releases page for what’s new (workspace epic #537, aethis-mcp#61). Non-blocking — the check runs in the background and never delays server startup — and fail-silent on any network error or timeout. Opt out with AETHIS_DISABLE_UPDATE_CHECK=1 (mirrors the same variable in aethis-cli); also skipped automatically when CI is set.
  • CI: cut a GitHub Release on publish. publish.yml now creates a GitHub Release for each published tag, using that version’s CHANGELOG section as the notes body (idempotent create-or-skip). Introduces the Releases channel on this repo — the subscribe-able “watch → releases” channel for the unified developer changelog (workspace epic #526, aethis-mcp#59). CI-only; no runtime or package change on its own.
aethis-cli
2026-07-20
  • feat: “what’s new” on aethis update. aethis update / aethis update --check now shows the changelog entries between your installed version and the latest release (titles + notes, newest ≤5, long notes truncated), sourced from the project’s GitHub Releases. If the Releases API is unreachable, rate-limited, or has nothing in range, it falls back to a link to the Releases page — the command never errors or hangs on this. The exit-time update banner also gained a “what’s new →” link to the same page. Coverage is forward-fill: only releases cut from here on populate the range, so an old install may see a gap. New update_check._fetch_github_releases(); the display logic lives in update_cmd._releases_in_range() / _print_whats_new(). (epic aethis-workspace#537)
  • ci: cut a GitHub Release on publish. The publish workflow now creates a GitHub Release for each just-published tag, using that version’s CHANGELOG.md section as the release notes, so the “watch → releases” subscribe channel stays current automatically. Idempotent (create-or-skip on an existing Release) and --verify-tag (never mints a synthetic tag). No package/runtime change. (epic aethis-workspace#526)
aethis-mcp
2026-07-19
Adds the Authoring Coach surface to MCP (aethis-mcp#57, workspace epic #514) — skill-building feedback for rule authors, advisory only, never a gate.Engine gate: the POST /api/v1/public/projects/{id}/review endpoint and the ambient review_hint fields are produced by aethis-core (epic phases P1/P4). This release must not be published to npm until that endpoint is live on api.aethis.ai; a released client calling a not-yet-deployed route would 404.
  • aethis_review_project (new tool). Reviews an authoring project against the deterministic authoring-coach rubric and renders the report: a score, per-check evidence across grounding / process / lifecycle, strengths, and the single highest-leverage next skill. Advisory only — it never blocks publishing. The deterministic layer needs no LLM key; coach=true (with an Anthropic key, via the usual anthropic_key_env / anthropic_key_keychain / anthropic_key forms) adds an opt-in LLM-synthesised coaching narrative on top. All server free-text (evidence / strengths / next-skill message / coaching) is fenced with fenceUntrusted before it reaches the model.
  • Ambient review_hint render. aethis_generate_and_test, aethis_refine, and aethis_publish now render a one-line coach hint when the server includes one on the response. The hint is computed entirely server-side (aethis-core P4); the client only renders it (fenced), never computes it.
  • X-Aethis-Client: mcp/<version> on every request. The client now sends a per-surface identifier header so the engine can attribute telemetry (e.g. review_hint-shown counts) to MCP vs CLI vs SDK.
  • 31 tools, up from 30. tests/tool-endpoint-map.ts and the drift suite are updated in the same change. Note: the drift suite’s live-alignment checks stay red against staging until the /review endpoint deploys there (expected epic ordering); the offline structural checks pass.
  • Tests. New mocked unit coverage in tests/client.test.ts (reviewProject request shape, the client-id header) and tests/server.test.ts (aethis_review_project render + fencing, coach key resolution, ambient hint render on generate/publish).
aethis-cli
2026-07-19
  • feat: aethis review [<project>] — the Authoring Coach report for a project. Runs the server-side rubric and prints an authoring score, 2–3 evidence-cited strengths, and the single highest-leverage next improvement (with its docs link and the lever that fixes it). Defaults to the current project in .aethis/state.json; pass a proj_… id to review any of your projects from anywhere. --verbose shows the full per-check table; --json (and any piped/--output json invocation) emits the raw ReviewReport. The deterministic report needs only your API key; --coach opts into LLM mentoring prose billed to your own Anthropic key (ANTHROPIC_API_KEY). Advisory only — the exit code is always 0 regardless of score. New AethisClient.review().
  • feat: every request now sends X-Aethis-Client: cli/<version> so the server can attribute per-surface telemetry (CLI vs MCP). The header carries no credentials and no PII, and is set once at client construction for all commands.
  • Requires aethis-core with the /api/v1/public/projects/{id}/review endpoint live (epic aethis-workspace#514, P1). The public release of this version is held until that endpoint is live on api.aethis.ai.
aethis-sdk-python
2026-07-17
  • feat(models): robot_hints + engine_version on the rulebook schema; engine_version on the ruleset schema. New RulebookSchemaResponse model (rulebook_id, sections, fields, robot_hints, engine_version) for GET /api/v1/public/rulebooks/{id}/schema — robot_hints is the rulebook’s natural-language conversational-agent guidance keyed by beat (general_context, preamble, session_start, postamble, session_end, stuck), None for a rulebook authored before the field existed. SchemaResponse (ruleset schema) gains engine_version: str | None = None for parity, also back-compat (defaults None when the engine doesn’t send it — true of the ruleset schema route today).
  • feat(models): graph/GraphResponse for the new /graph endpoint. New GraphResponse (ruleset_id/rulebook_id, slug, name, graph, mermaid) and RulesetGraph (nodes, edges, sections, stats) model the ruleset/rulebook dependency graph (field → criterion → group → outcome) plus its rendered Mermaid diagram. Node/edge shape varies by node type, so nodes/edges stay loosely-typed dicts rather than a rigid per-type schema — deliberately permissive so a legacy or empty graph (nodes: []) still parses.
  • feat(client): get_graph(ruleset_id) (sync + async) — wraps GET /api/v1/public/rulesets/{id}/graph, returning GraphResponse. Public rulesets can be inspected without an API key, same as get_schema().
  • feat(decide): include_graph_overlay parameter on decide() / decide_rulebook() (sync + async), and a matching graph_overlay: dict[str, Any] | None = None field on DecideResponse. Set include_graph_overlay=True to get this decision’s per-criterion status stamped onto the ruleset’s dependency graph, in the same shape get_graph() returns.
  • All additions are additive and backwards-compatible: every new field defaults to None/False/an empty collection, so a legacy response (no robot_hints, no engine_version, no graph_overlay) still deserialises unchanged.
aethis-mcp
2026-07-17
Propagates the aethis-core 0.37–0.40 authoring batch to the MCP surface (aethis-mcp#49, workspace epic #327). Engine gate: live on api.aethis.ai 0.45.2, confirmed via the drift suite’s live-alignment checks before this release.
  • aethis_graph (new tool). Fetches the ruleset-map graph — either for a single published ruleset (ruleset_id, may be public/anonymous for a public showcase ruleset) or a composed rulebook (rulebook_id, always requires an API key) — the same mutual-exclusivity shape as aethis_decide. Returns {ruleset_id|rulebook_id, slug, name, graph: {nodes, edges, sections, stats}, mermaid}: each node’s display.sentence/display.routes/ display.expr shows how that branch composes, and mermaid is a ready-to-render diagram string.
  • include_graph_overlay on aethis_decide (additive). Stamp a specific decision’s per-criterion outcome (satisfied/not_satisfied/pending) onto that same graph and return it as graph_overlay in the decide response — a “you are here” map for those inputs. Off by default; the response is unchanged when omitted.
  • aethis_create_rulebook / aethis_update_rulebook (new tools). Create an empty draft Rulebook (name/domain/slug/description) or update one, both accepting robot_hints — beat-keyed natural-language guidance for the conversational agent (active beats: general_context, preamble, session_start, postamble, session_end, stuck; reserved: persona, conversational_style, section_transition). An unknown beat is rejected client-side before the round-trip, mirroring aethis-cli’s _validate_robot_hints (v0.23.0). Rulebook composition (outcome_logic, ruleset_refs) is a larger surface not covered by these two tools yet.
  • years_between in the DSL helper reference (README). Documents the new completed-whole-years, leap-correct date operator (mirrors aethis-core Operator.YEARS_BETWEEN, commit 3607558) alongside days_between so generation can use it for age-from-date-of-birth instead of days_between(...) / 365 (division isn’t supported anyway).
  • 30 tools, up from 27. tests/tool-endpoint-map.ts and the drift suite are updated in the same change; every new operation/field/param is verified against the live api.aethis.ai OpenAPI document (engine 0.45.2).
  • Tests. New mocked unit coverage in tests/client.test.ts / tests/server.test.ts for the graph client methods, the create/update rulebook client methods, robot_hints beat validation (known + unknown + reserved), and include_graph_overlay pass-through; the nightly staging integration lane gains a real aethis_graph fetch, a aethis_decide include_graph_overlay round-trip, and an aethis_create_rulebook → aethis_update_rulebook robot_hints round-trip (with best-effort archive cleanup of the probe rulebook).
aethis-cli
2026-07-17
  • feat(rulebooks): aethis rulebooks graph <id> — fetch and render the rulebook-level ruleset-map dependency graph (field -> criterion -> group -> outcome). Prints a node-count summary + a table of nodes (id, type, the criterion’s human-readable display.sentence, field count); --mermaid prints the raw Mermaid diagram source for piping into a renderer; --output json returns the full payload ({rulebook_id, graph: {nodes, edges, sections, stats}, mermaid}), including each node’s display.routes/display.expr for programmatic consumers. This endpoint requires a valid API key even for a public rulebook (confirmed against the live engine) — unlike the ruleset-level graph below, there’s no anonymous path. New AethisClient.get_rulebook_graph().
  • feat(rulesets): aethis rulesets graph <ruleset_id> — the single-ruleset analogue, open for public rulesets with no API key required (load_client_or_anon). Same table/--mermaid/--output json shape. New AethisClient.get_ruleset_graph().
  • feat: --include-graph-overlay on aethis decide and aethis rulebooks decide — stamps the decision’s per-criterion status onto the rule-map graph, returned as a graph_overlay field on the response (--output json to inspect it). Additive request flag; a plain-text hint is printed when the overlay is present and JSON wasn’t explicitly requested.
  • feat(rulebooks): aethis rulebooks schema surfaces engine_version. The schema response already carries the aethis-core build that served it (e.g. aethis-core@0.45.2); the CLI now prints it as a header line ahead of the schema payload instead of leaving it buried in the JSON.
  • Requires aethis-core 0.40.0+ (live on api.aethis.ai / staging.api.aethis.ai as of this release) for /graph, include_graph_overlay, and engine_version on /schema. robot_hints (shipped v0.23.0) is unaffected by this release.
aethis-core
2026-07-16

Fixed

  • An API key with expires_at set no longer 500s every request: the Mongo-stored (tz-naive) expiry is normalized to UTC before comparison, so an expired key gets its intended 401 api_key_expired. Latent since the expiry field existed — no key carried a non-None expiry until 2026-07-16 (issue #275).
aethis-sdk-python
2026-07-15
  • feat(errors): typed 401/403/429 exceptions carrying the structured error envelope. classify_response now raises AethisAuthError (401), AethisPermissionError (403), or AethisRateLimitError (429) — each a subclass of AethisAPIError, so existing except AethisAPIError handlers keep catching them (non-breaking). AethisError gains .reason_code, .missing_permissions, and .hint, lifted out of the public API’s structured envelope ({"detail": {"error", "reason_code", "missing_permissions", "hint", ...}}), so a caller can branch on err.reason_code == "denied_missing_permission" or read err.missing_permissions without re-parsing err.body. Plain-string and FastAPI-422-list details are untouched (fields stay None / []). Constructor stays backwards-compatible (new args default to None).
  • test(staging): live integration lane against staging.api.aethis.ai. New tests/integration/ (marker staging, excluded from the PR gate) mints a real API key the way a user does — Clerk sign-in ticket → frontend-API JWT → POST /api/v1/keys/ → teardown — and exercises every public method on Aethis + AsyncAethis (decide, decide_rulebook, list_rulesets, get_schema, whoami, explain, explain_failure, get_source, sync/async session flows) plus live 401/403 typed-error assertions and a contract cross-check. Reports red (never green-by-skip) when creds are missing or staging/contract is unreachable.
  • test(parity): recorded-live fixture parity. tests/shapes.compare_shape diffs the mocked conftest fixture builders (make_decide_response, make_schema_response, make_ruleset_summary) against real staging payloads so the mocked suite can’t silently drift from reality; the builders were updated to match the current engine shape (slug/rulebook_id, graph_overlay/timing, richer next_question/schema fields).
  • chore(ci): coverage floor (--cov-fail-under=45) + staging marker + nightly staging-integration.yml (report-only, workflow_dispatch + schedule, uploads a qa-run-record artifact for the sdk-staging lane). The coverage flags live in the CI command, not in addopts, so a bare pytest / uv run pytest works without pytest-cov (which is only in the dev extra) installed; the floor is still enforced in CI.
aethis-mcp
2026-07-15
Test-infra only — no runtime/behaviour change to the server or its tools.
  • Tool-schema drift suite (tests/drift.test.ts). Guards that the 27 server.tool() input schemas never silently drift from the engine. Reads each tool’s real zod shape (no vendored schema copy) and compares field names, types, and required-ness against the deployed staging OpenAPI document — the oracle. An explicit, checked-in tests/tool-endpoint-map.ts records the tool → operation correspondence and field renames (e.g. force → force_unsafe); a tool missing from the map, an unknown extra tool, an unclassified input field, a mapped operation absent from the engine, or a mapped body field the engine no longer has all fail loud. Runs in the PR gate (network-tolerant) and nightly (network-required).
  • Staging integration lane (tests/integration/, nightly). Runs the built server as a subprocess with a freshly minted staging key and drives it over the real MCP protocol: tools/list (== 27), a read-only core loop, and aethis_decide against a public showcase ruleset; a negative path proves an invalid key returns a structured error result while the server stays alive. Keys are minted via the self-serve path (server-default scopes), named e2e-dx-mcp-*, and revoked + swept in teardown.
  • staging-integration.yml — nightly + manual, report-only, emits a QA-run-shaped run record artifact for downstream ingestion; missing secrets or unreachable staging fail red, never skip-green.
aethis-cli
2026-07-15
  • feat: authorization errors now render the server’s hint and the missing scope readably. A 403 denied_missing_permission (and 401) previously printed the raw error object on commands that render their own errors (projects, whoami, …); the CLI now renders one clean line naming the missing permission plus, on its own dim line, the server’s follow-up hint (e.g. how to request access). The top-level handler and the per-command renderer now share one formatter (aethis_cli.output.format_error_detail / render_api_error), so every command surfaces the same readable message. The hint is rendered with markup disabled (so a hint containing [brackets] isn’t dropped) and non-string missing_permissions items are coerced (so a server quirk can’t turn the error into a traceback).
  • test: new staging integration lane (tests/integration/, marker staging). Acquires an API key the self-serve way (a fenced e2e user’s session → mint with the server’s default scopes, no scopes field), drives the CLI core loop against deployed staging (whoami/status, projects list/archive, rulesets/explain/fields/decide against a public showcase ruleset), and asserts the negative paths a caller actually sees — a scope-reduced key’s 403 and a revoked key’s 401 — with the error envelopes checked against the machine-readable public-API contract. Report-only nightly workflow (staging-integration.yml); never gates a merge. Run locally with the one-liner in tests/integration/README.md.
  • test: the spacecraft authoring e2e moved to its own weekly lane (authoring-e2e-weekly.yml). It drives the LLM authoring pipeline, so it is kept out of the nightly LLM-free cadence; the model is passed explicitly via X-Anthropic-Key, generation is bounded by an explicit iteration cap (SPACECRAFT_GENERATION_TIMEOUT), and the manual marker stays as the local escape hatch.
aethis-core
2026-07-15

Added

  • Read-only cross-tenant /api/v1/admin/* router (epic aethis-workspace#480, P1): /admin/keys, /admin/usage (daily rate-limit counters), /admin/generation-jobs (+ /{id} with trace), /admin/decisions (list excludes field_values; /{decision_id} full record), /admin/publish-audits, /admin/rulebooks, /admin/rulesets (metadata; DSL source additionally requires rulesets:source). Gated by the new internal-only admin:read scope AND api_key.internal == true AND a hard-fail under DISABLE_AUTH=true (503) — scope alone is deliberately not the boundary. admin:read is never self-servable (ALLOWED_SCOPES) and never an alias target. Uniform {items, next_cursor, limit, skipped_invalid} envelope, cursor pagination, server-side max page size, default 30-day lookback on time-series lists, per-doc validation (one malformed record never 500s a list), typed filters (tenant_id=None only via explicit anonymous_only=true, never a wildcard). New admin rate-limit category enumerated in every tier (internal ≈ unlimited). Routes stay in the live /openapi.json but out of the published public contract.
aethis-core
2026-07-15

Added

  • Decision log (epic aethis-workspace#480, P0): every /decide call can now be persisted server-side as a DecisionRecord (collection decisions), fulfilling the DecideResponse docstring’s deferred “server-side audit persistence”. Gated by DECISION_LOG_ENABLED (default off, fail-closed) with TTL retention via DECISION_LOG_TTL_DAYS (default 90) — the TTL index is declared in code and asserted at boot (logging disables loudly if the live index lacks expireAfterSeconds). The write is fire-and-forget off the hot path: bounded pending-task set, client-side insert timeout, fully exception-isolated (no decision-log failure can alter a /decide response). Write failures log structured ERROR, increment a counter, and can alert via DECISION_LOG_ALERT_WEBHOOK (Google Chat, rate-limited; unset = off).
  • DecideRequest.caller_ref — optional opaque caller metadata (flat string→string dict, ≤2 KB, no $/dotted keys) stored on the decision record for the caller’s own attribution (e.g. tda-server sends {firm, application_id}). Never an authorization key or cross-principal predicate (defect shape DS-25); invalid values are dropped with a warning, never rejected.
aethis-core
2026-07-10

Added

  • FieldDefinition (and the /schema + /decide next_question / optimal_path envelope) now carries an optional x_ui_widget: Optional[str] authoring override. Currently only "free_text" is recognised: it tells downstream consumers (Lisa’s expected_input emission in tda-server) to suppress the schema-derived structured-answer affordance (chips / select / date-picker) for that field and render a plain text composer, even though the field has a typed sort. Defaults to None — purely additive. (#255, epic aethis-workspace#422)
aethis-core
2026-07-10

Added

  • /decide: the next_question (and each optimal_path entry) now carries optional sort and enum_values fields, exposing the field’s answer type (Int / Bool / String / Enum / Date / Duration) and, for Enum sorts, its allowed values. Lets callers render typed input affordances (yes/no chips, a date picker, an option list) without a second /schema round-trip. Both default to None, so the change is additive — existing consumers are unaffected. Mirrors the sort / enum_values pair already on /schema’s FieldInfo. Populated on the ruleset path; rulebook /decide callers continue to source the type from /schema. (#254, epic aethis-workspace#422)
aethis-mcp
2026-07-08
  • docs: correct stale paper citation in the construction-insurance demo. The demo cited the withdrawn v3.6/v3.7 claim that GPT-5.4 at reasoning_effort=low scores 7/11 on the exception-chain subset; the paper withdrew that result in v3.8 (instrumented replication: 11/11). The demo now attributes 7/11 to GPT-5.3 only and pins the paper citation at v3.11. No code changes.
aethis-sdk-python
2026-07-04
  • fix(errors): attach the API’s detail (and full body) to AethisAPIError. On the primary error path, classify_response parsed the 4xx detail only to log it, then raised AethisAPIError("Aethis API returned 422") — blinding callers to why the request failed. The exception message now reads "Aethis API returned 422: <detail>" when a detail is present, and AethisError gained .detail / .body attributes carrying the parsed payload (both None for timeouts / connection errors). Constructor signatures stay backwards-compatible (new args default to None).
  • fix(models): DecideResponse.explanation is a single object, not a list. The field was typed list[dict] | None but the engine returns Optional[Dict[str, Any]] ({decision, groups: [...], unused_facts: [...]}), a latent ValidationError for any caller that actually requested one. Retyped to dict[str, Any] | None.
  • feat(decide): include_explanation parameter on decide() / decide_rulebook() (sync + async). The engine has always accepted include_explanation on POST /decide, but the SDK never sent it, leaving DecideResponse.explanation permanently None. Passed through in the request payload alongside include_trace; defaults to False.
  • feat(models): typed FieldNote and NextQuestion.notes. The engine attaches structured author guidance (note_text, source, metadata) to each next_question; the SDK silently dropped it. Adds the FieldNote model (exported from the package) and notes: list[FieldNote] on NextQuestion, defaulting to [] so older responses without notes keep parsing.
  • feat(client): list_rulesets(limit=20, offset=0) (sync + async) — wraps GET /api/v1/public/rulesets, returning the previously-exported-but-unreachable RulesetSummary model. Anonymous callers get public rulesets; an API key additionally surfaces that key’s own rulesets. limit is clamped by the engine to 1-50.
  • docs(readme, _base): capability-table + docstring fixes. README’s “What’s included” table now lists explain_failure, decide_rulebook, list_rulesets, include_explanation, and FieldNote; the build_headers docstring no longer names a nonexistent /next_question endpoint.
aethis-mcp
2026-07-04
Cross-surface review batch (aethis-mcp#50).
  • Surface next_question.notes (additive). aethis_next_question now renders a Notes block after the question when the ruleset author attached notes to it (each note carries note_text, source, and metadata). Notes are labelled by metadata.type (e.g. why, legal_background) when present, and each note’s text is wrapped with fenceUntrusted(...) since it is author-provided server content. Output is unchanged when no notes are present. The tool description now mentions the Notes block.
  • Fence aethis_list_guidance output. The guidance_text and source fields returned by the server were interpolated into the tool result unfenced, unlike every sibling handler. They are now wrapped with fenceUntrusted(...) under the UNTRUSTED_PREFACE warning (GHSA-ph7q-r9q4-922g hardening).
  • Send a single provider header. The per-request LLM key was sent under both X-Anthropic-Key and X-OpenAI-Key. It is now sent only as X-Anthropic-Key, matching how resolveLlmKey resolves the key.
  • Correct the stale latency figure. Two guidance strings claimed decisions are <5ms; corrected to <1ms to match the README and the canonical figure.
  • Docs: rewrote the CLAUDE.md architecture section to the real layout (src/index.ts + src/client.ts + src/credentials.ts, tests under tests/) instead of the non-existent src/server.ts + src/tools/ tree.
aethis-cli
2026-07-04
  • fix: network errors now render one actionable line, not a raw traceback. When the API is unreachable, times out, or a DNS/TLS error occurs, every command now prints Could not reach the Aethis API at <url>: <reason>. plus a “check your connection” hint and exits non-zero, instead of dumping an httpx stack trace. The top-level handler catches httpx.HTTPError (the umbrella over connect/timeout RequestErrors), matching the graceful handling login/account already had.
  • feat: non-interactive environments bypass confirmation prompts. A truthy AETHIS_NONINTERACTIVE or CI env var (values 1/true/yes, case-insensitive) now flips the whole process non-interactive, so destructive commands (account revoke, rulesets archive, projects archive, rulebooks archive, rulebooks tests delete) proceed without waiting on stdin, so a background job or CI step no longer hangs on a [y/N] prompt. The bypass prints a one-line notice so it’s never silently active. The explicit per-command --yes/-y flags keep working unchanged. New shared aethis_cli.prompts.confirm_or_abort helper.
  • docs: refreshed the worked examples in decide/explain/fields help to use public showcase rulesets (aethis/spacecraft-crew-certification, aethis/consumer-credit-prequalification) instead of product-specific slugs.
  • chore: make install uses uv pip install -e ".[dev]" (matching the README) instead of bare pip.
  • minor: dropped the unused upgrade-command strings from update_check._detect_install_method (it now returns just the detected method; the concrete upgrade argv is still built by update’s _upgrade_argv); decide reads the decision field with a safe default so a payload without decision renders as unknown rather than raising KeyError.
aethis-cli
2026-06-25
  • feat(rulebooks): declare robot_hints: in a rulebook file and push them to the engine. Rulebook authors can now provide natural-language guidance for the conversational assistant alongside the rulebook’s other configuration.
    • aethis rulebooks create <name> --file rulebook.yaml — a new --file/-f option reads a robot_hints: block (a sibling of name/domain/outcome_logic) from a rulebook.yaml/.json and sends it on create. CLI flags still own name/domain/slug/description; only the hints are taken from the file. No --file (or a file without a robot_hints: key) is a clean no-op — behaviour is unchanged.
    • aethis rulebooks set-logic <id> -f rulebook.yaml now also accepts a wrapped form: when the top-level object carries an outcome_logic: key, a sibling robot_hints: block is pushed in the same update. A bare Expr AST file (the prior shape) is still accepted unchanged.
    • robot_hints is a mapping of beat-name to a natural-language string. Active beats: general_context, preamble, session_start, postamble, session_end, stuck. Reserved beats (accepted, not yet acted on): persona, conversational_style, section_transition. Unknown beat keys and non-string values are rejected client-side with a clear message before the round-trip.
    • New optional robot_hints parameter on AethisClient.create_rulebook() / update_rulebook(); omitted from the request body when not supplied, so calls against an older engine are unaffected.
    • Requires aethis-core with the rulebook robot_hints field (aethis-core#220); mid-deploy to staging at time of writing. Against an engine without it, the field is ignored/rejected server-side.
aethis-cli
2026-06-16
  • feat(fields): aethis fields is now a command group for the full field-authoring loop. Bare aethis fields [-b <ruleset>] still shows a ruleset’s field schema (unchanged); three subcommands manage the local fields/fields.yaml:
    • aethis fields discover — uploads the project’s sources/ (creating the project if needed), runs server-side LLM field discovery, and merges the proposals into fields/fields.yaml so you start from a real draft instead of a blank file. Existing entries are preserved — only new keys are appended — so hand-authored labels/questions/hints are never clobbered. Prints the completeness score and any critical gaps. Needs an LLM key (ANTHROPIC_API_KEY), same as generate; without one it fails with a clear message naming the env var instead of a raw server header error. New AethisClient.discover_fields().
    • aethis fields pull — syncs the server’s authoritative produced fields (key + type + enum values) back into fields/fields.yaml so local matches reality after a generate. Local-only label/hints are preserved; fields absent from the server schema are kept and reported rather than silently dropped.
    • aethis fields validate — checks fields/fields.yaml before upload: valid type (int/bool/string/enum/date/duration), no duplicate keys, enum requires enum_values. The same validation now also runs inside aethis generate, per contributing file (rulebook + ruleset), so duplicate keys within a file fail fast before any server state changes.
    • discover/pull only ever write a vocabulary that re-validates: an unknown server type or an enum with no values falls back to string instead of producing a file the next validate/generate would reject. Writes also preserve any hand-authored keys the tool doesn’t model (e.g. description) rather than dropping them.
  • feat(generate): the field spec/produced diff is surfaced after generation. After a successful aethis generate, the CLI compares the pinned field vocabulary against what the engine actually produced and prints pinned-but-not-produced / produced-but-not-pinned fields (with a pointer to aethis fields pull) instead of the drift passing silently.
  • feat(init): rulesets can declare rulebook membership explicitly. A rulebook: key in a ruleset’s aethis.yaml (a path to the enclosing rulebook) now declares membership directly; the directory-position convention (<rulebook>/rulesets/<ruleset>/) remains the fallback. The init scaffold documents the key.
  • perf: source uploads are now idempotent. discover and generate share one project-resolution + upload path, and a per-file mtime ledger in .aethis/state.json means a discover followed by a generate (or repeated generates) only re-uploads sources that actually changed instead of re-pushing the whole sources/ tree each time.
  • fix(generate): don’t lose the ruleset id on a fast success. The poll loop occasionally saw the job flip to success a beat before latest_ruleset_id was populated, writing a null id to state and leaving fields pull / the field diff with nothing to work from. It now re-polls briefly for the id and only records a real one — never clobbering a prior good id with null.
  • example + e2e: examples/community-grants-rulebook/ is a generic rulebook (one shared field) with two member rulesets, and tests/e2e/test_rulebook_hierarchy_e2e.py (gated by the manual marker) drives discover/validate/generate/pull against a live API and asserts the shared rulebook field propagates into both members.
  • No engine change required — all endpoints (/fields/discover, /rulesets/{id}/schema, /fields/spec) are already served by aethis-core and used by the MCP server.
aethis-cli
2026-06-16
  • feat(init): field definitions get a real home (fields/fields.yaml). aethis init now scaffolds a fields/ directory with a fields.yaml for declaring the field vocabulary (key + type + optional label/question/hints). Previously fields had no dedicated home and only surfaced implicitly as the inputs: keys inside tests/scenarios.yaml. aethis generate reads fields/fields.yaml, pins the field keys/types via the project field-spec endpoint, and routes each field’s label/question/hints through guidance so a field is defined once.
  • feat(init): --kind rulebook scaffolds a rulebook. aethis init <name> --kind rulebook lays down a rulebook directory with shared guidance/ and fields/ plus a rulesets/ directory for member rulesets. When a ruleset lives under a rulebook (<rulebook>/rulesets/<ruleset>/), aethis generate propagates the rulebook’s guidance hints and field vocabulary into the ruleset — the rulebook definition wins on shared field keys — so a common field (e.g. date of birth) is defined once at the rulebook level and the end user is asked for it only once. --kind defaults to ruleset, so existing behaviour is unchanged.
    • New AethisClient.set_field_spec() (project field-spec endpoint, already served by aethis-core / used by the MCP server). No engine change required.
aethis-cli
2026-06-03
  • feat(rulebooks list): anonymous fallthrough to the public rulebook catalogue. With no cached API key, aethis rulebooks list now lists the cross-tenant public catalogue (rulebooks with public visibility, active status) instead of printing the v0.19.1 pointer message — completing the parity with aethis rulesets list. A dim one-liner (“No API key — showing public rulebooks…”) distinguishes the anonymous view; with a key, the tenant listing is unchanged.
    • New AethisClient.list_public_rulebooks(); use with make_anonymous_client so a cached key doesn’t promote the call to an authenticated tenant listing.
    • Requires aethis-core v0.29.0+ on the target API (live on api.aethis.ai). Against an older engine the anonymous path surfaces the server’s 401 cleanly.
aethis-cli
2026-06-03
  • fix(rulebooks list): stop prompting browser sign-in for anonymous users. aethis rulebooks list with no cached API key used to trigger the lazy-auth browser login — bad first-contact DX for a read-only browse command. Rulebooks are tenant-scoped, so an anonymous caller has nothing to list; the command now prints a pointer to the anonymous public catalogue (aethis rulesets list) and to aethis login, and exits 1 without ever opening a browser.
    • True anonymous fallthrough (listing public rulebooks without an account, mirroring aethis rulesets list) needs engine support for a public rulebook catalogue and is tracked separately; this release removes the login prompt in the meantime.
aethis-cli
2026-06-03
  • feat(update): aethis update — self-update the CLI to the latest release. Detects how the CLI was installed (uv tool, pipx, or pip) and runs the matching upgrade command. aethis update --check reports whether a newer release exists without installing anything.
    • Editable (development) installs are refused with a pointer to git pull && uv sync instead of clobbering the checkout.
    • The exit-time “new release available” banner now points at aethis update rather than a method-specific command.
    • fix: the banner’s uv upgrade hint was uv tool install --upgrade aethis-cli, which re-resolves from scratch and silently drops any extra --with requirements (e.g. plugin packages installed alongside the CLI). Both the banner’s install-method detection and aethis update now use uv tool upgrade aethis-cli, which honours the original install receipt.
    • A successful (or no-op) aethis update refreshes the banner’s 24h cache, so the notice goes quiet immediately after updating.
aethis-mcp
2026-05-29
aethis_refine now performs finding-driven incremental re-authoring: it seeds generation from the section’s active ruleset and asks the engine for the minimal edit to fix failing test cases while keeping passing tests green, instead of re-authoring the whole section from scratch. aethis_generate_and_test is unchanged (from-scratch authoring).Why this matters: fixing one wrong case in a published ruleset previously meant a full-section regenerate — expensive, and prone to silently regressing carefully tuned behaviour (e.g. caseworker-review criteria that intentionally yield undetermined). Refine keeps the blast radius to the criteria that actually need to change; the full-suite gate still guarantees no regression ships.Requires aethis-core with the mode parameter on /generate (engine ≥ the release shipping seed-from-existing refine). Older engines ignore the body and fall back to from-scratch generation.
  • client.generate() / generateAndTest() accept an optional mode and send {mode:"refine"} on the generation request body.
aethis-cli
2026-05-29
  • feat(refine): aethis refine + aethis generate --mode refine for incremental, seed-from-existing re-authoring. Instead of re-authoring a whole section from scratch, refine seeds generation from the section’s active ruleset and makes the minimal edit to fix failing tests while keeping passing tests green.
    • aethis refine [--hint "..."] [--seed-ruleset-id <id>] — the phase-3 TDD-loop command: optionally add a guidance hint, then refine. Defaults to seeding from the section’s active ruleset.
    • aethis generate --mode refine [--seed-ruleset-id <id>] — the same capability via a flag on generate; --mode fresh (default) is unchanged from-scratch authoring.
    • AethisClient.generate() gains optional mode / seed_ruleset_id; a no-arg call still sends no body, so it stays backwards-compatible against engines without the parameter.
    • Requires aethis-core with the mode parameter on /generate (live on api.aethis.ai). Against an older engine the flags no-op (empty body = fresh).
aethis-mcp
2026-05-27
Add the rulebook tier to the MCP read surface. Closes aethis-mcp#43 for the two endpoints the engine exposes today; the public-catalogue equivalent (aethis_discover_rulebooks) is deferred until aethis-core ships a no-auth rulebooks catalogue endpoint.Why this matters: until now, an agent connected via MCP could see the parts (rulesets via aethis_discover_rulesets / aethis_list_rulesets) and evaluate the whole (aethis_decide with rulebook_id), but had no way to find a rulebook or inspect how its rulesets compose. Concrete failure mode from a real 2026-05-27 session: asked whether aethis/uk-fsm was “three rulebooks or one rulebook with three rulesets”, the MCP gave no read path that could answer.

Added

  • aethis_list_rulebooks — lists rulebooks in the current tenant (auth-required, tenant-scoped). Returns the fields needed to distinguish one composed rulebook from N independent rulesets: rulebook_id, slug, name, domain, status, version, outcome_logic (the composition Expr AST), ruleset_refs, timestamps. Mirrors aethis_list_rulesets.
  • aethis_rulebook_schema — fetches one rulebook’s composition, bridged rulesets (with names + slugs + ruleset_ids), and aggregated input fields. Accepts either a slug (aethis/uk-fsm) or an opaque rb_* id. Mirrors aethis_schema but at the rulebook tier.
  • AethisClient.listRulebooks() / getRulebookSchema() — wrap GET /api/v1/public/rulebooks/ and GET /api/v1/public/rulebooks/{slug-or-id}/schema. The schema helper preserves the literal / in slugs (so aethis/uk-fsm hits the engine’s {namespace}/{name} matcher) and URL-encodes opaque ids.

Deferred

  • aethis_discover_rulebooks — the cross-tenant public catalogue equivalent of aethis_discover_rulesets. The engine’s /api/v1/public/rulebooks/ endpoint is tenant-scoped + auth-required on prod today; no anonymous catalogue variant exists. Will land once aethis-core adds it.
aethis-cli
2026-05-27
  • feat(output): gh-style machine-readable output mode (--output json, --json fields, --jq). Every list/show command (and the decision commands) now emit structured JSON on demand, so aethis rulesets list --output json | jq '.[0].slug' just works instead of trying to scrape ANSI-coloured Rich tables.
    • --output table|json — pick the format. Default: table on a TTY, json when piped (matches gh’s pipe-friendly autodetect).
    • --json FIELDS — implies --output json; takes a required comma-separated value (--json id,name) that limits the payload to those fields. (gh’s bare---json introspection trick is not yet exposed — Click/Typer’s option parser can’t cleanly distinguish “flag with no value” from “flag followed by positional”, so it’s deferred to a future --list-fields flag.)
    • --jq EXPR — pipe JSON output through jq before printing. Requires the jq binary on PATH; clear error with install hint if missing.
    • Commands migrated: rulesets list/show, rulebooks list/show/get-fields/tests list/schema/explain/decide, projects list/show, account keys, profile list, guidance list, fields, explain, decide, status. Each command has a sensible JSON shape — status --output json | jq .identity.key_id returns the live key id without rooting through any prose.
    • Footer hints (Try: aethis ...) are suppressed in JSON mode so pipes get clean output.
    • New module aethis_cli/render.py is the single emit point; new test file tests/test_render.py covers the matrix.
  • breaking(guidance export): --output renamed to --output-file to avoid clashing with the new global --output flag. Short form -o unchanged. Affects scripts that pipe to a named file: aethis guidance export --output foo.yaml → aethis guidance export --output-file foo.yaml (or -o foo.yaml).
aethis-cli
2026-05-27
  • fix(status, whoami): read the multi-profile credentials file the same way every other command does. aethis login --api-key ... writes profiles.<name>.api_key to ~/.config/aethis/credentials (the multi-profile schema introduced in v0.10), but aethis status and aethis whoami had stale local resolvers that only looked for a flat top-level api_key (and whoami was looking at the wrong filename, credentials.yaml). Result: after a fresh aethis login, aethis status reported no API key and aethis whoami reported No Aethis API key configured, even though the same key worked for aethis projects list, aethis generate, and every other authoring command.
    • Both commands now route through the canonical resolve_cached_key() helper in auth_helpers.py, which honours AETHIS_API_KEY env → active profile → keychain → legacy .yaml file.
    • The _resolve_cached_key symbol is renamed to resolve_cached_key (public). The legacy _resolve_key_silent (status_cmd) and _resolve_api_key_lax (whoami_cmd) are removed.
    • Regression test in tests/test_status_cmd.py writes a real multi-profile credentials YAML to a temp XDG_CONFIG_HOME and asserts both commands surface the key.
aethis-mcp
2026-05-26
  • chore(server): tighten MCP instructions against decision extrapolation. Adds a “Reporting decisions” section to the server instructions block (visible to every client model as part of its system prompt on connect). New rules forbid asserting facts that are not in the tool response, generalising a single-ruleset decision to a composite outcome, naming rulesets/rulebooks not yet observed in the session, and offering follow-up calls against unverified slugs. Triggered by a real user trace where a model summarising a uk-fsm-child-eligibility decision closed with an offer to run the broader aethis/uk-fsm rulebook “to get the complete household-level decision” — the rulebook exists but currently 422s on prod (empty ruleset_refs, see aethis-core#90), so the offer overstated what would actually happen. Advisory, not enforced — but client models reliably honour instructions blocks.
aethis-sdk-python
2026-05-25
  • feat(explain-failure): Aethis.explain_failure() + AsyncAethis.explain_failure() — wraps POST /api/v1/public/rulesets/{ruleset_id}/explain-failure, returning the failing criterion and a targeted fix hint for a mismatched /decide result. Accepts field_values, expected_outcome ("eligible" | "not_eligible" | "undetermined"), and an optional test_name (default "test"). Return type is dict[str, Any] to match explain() / get_source() — can be tightened once the response shape stabilises. Note: ruleset_id must be the concrete identifier (not a slug); the underlying endpoint does not currently resolve slugs. Previously, callers had to drop to raw httpx for this endpoint — flagged in recipes/evaluate-a-case.mdx and recipes/debug-a-decide.mdx.
aethis-sdk-python
2026-05-22
  • docs(readme): rulebook surface advertised on the PyPI landing page. The v0.5.0 release shipped decide_rulebook() and rulebook_id on DecideResponse, but the README still framed the SDK as ruleset-only. Adds a dedicated “Composed rulebook” section with a runnable UK FSM example, the always-scope-gated note, and the async equivalent.
  • docs(install): switch pip install to uv add per workspace no-pip rule. The PyPI landing page is a public-facing surface bound by .claude/rules/no-pip.md. Adds uv pip install as a venv-friendly alternative.
  • docs(engine_version): update sample audit-field comment from aethis-core@0.10.0 to aethis-core@0.27.0 — matches live prod engine.
  • docs(beta): clarify that decision endpoints are anonymous only for single rulesets — rulebook decide is always scope-gated, so the SDK’s “anonymous when no key” claim needed a footnote.
aethis-sdk-python
2026-05-22
  • feat(rulebook): Aethis.decide_rulebook() + AsyncAethis.decide_rulebook() — evaluate a composed multi-ruleset rulebook through the SDK. Mirrors decide() but sends rulebook_id in the payload. Accepts either an opaque rb_<id> or a slug (e.g. aethis/uk-fsm). Requires an API key — rulebook evaluation is always scope-gated. Closes #14. Requires aethis-core v0.27.0+ live on the target API for slug-form rulebook paths.
  • feat(models): add rulebook_id: Optional[str] to DecideResponse — surfaces the rulebook identifier when the response was a composed-rulebook decide. Backwards-compatible: ruleset-only decides keep rulebook_id=None.
aethis-mcp
2026-05-22
  • docs(readme): v0.27.0 accuracy pass. Three fixes for fresh-developer accuracy:
    • Documented rulebook_id as an alternative to ruleset_id on aethis_decide — mutually exclusive; composed-rulebook evaluation always requires an API key.
    • Quickstart example: corrected field name from species to space.crew.species (the actual field ID in the spacecraft-crew-certification ruleset).
    • Windsurf config path: corrected from .windsurf/mcp.json to ~/.codeium/windsurf/mcp_config.json (canonical path per aethis-cli README).
aethis-mcp
2026-05-22
Add rulebook surface to aethis_decide — closes the converged-2-term client-completeness gap for MCP. The tool now accepts either ruleset_id (single ruleset) or rulebook_id (composed rulebook), mutually exclusive. Mirrors aethis-sdk-python v0.5.0 and aethis-cli rulebooks decide.

Added

  • aethis_decide tool accepts rulebook_id as an alternative to ruleset_id. Pass an opaque rb_<id> or a slug like aethis/uk-fsm. Composed-rulebook evaluation is always scope-gated by the engine — anonymous callers get HTTP 401.
  • AethisClient.decideRulebook(rulebookId, fieldValues, options?) — parallel to decide(); sends rulebook_id in the /decide payload.

Changed

  • aethis_decide description and schema updated to reflect both paths. Tool validates that exactly one of ruleset_id / rulebook_id is provided.

Requires

  • aethis-core v0.27.0+ live on the target API for slug-form rulebook paths. The rulebook_id body field on /decide has been supported since aethis-core v0.18.x.
aethis-cli
2026-05-22
  • docs(readme): v0.27.0 accuracy pass. Three fixes for fresh-developer accuracy:
    • Install block: removed the pip install fallback (uv and pipx are the recommended forms per workspace policy). Development section: pip install -e ".[dev]" → uv pip install -e ".[dev]".
    • Added Rulebooks command-group section documenting the converged 2-term model surface shipped in v0.14.0–v0.16.1 (aethis rulebooks + aethis rulesets promote-to-live).
    • Updated engine_version example to aethis-core@0.27.0 (was absent; clarified to current production version).
aethis-cli
2026-05-22
  • docs(rulebooks set-logic): the docstring example for field_ref.key now matches engine behaviour. Phase A.16 (aethis-core v0.26.0+) added per-section aggregate group synthesis, so field_ref.key = <ruleset_name> resolves to the AND of that ruleset’s groups. The unscoped group-name and scoped <ruleset_name>.<group> forms remain available for advanced compositions. Requires aethis-core v0.26.0+ live on the target API.
aethis-cli
2026-05-21
  • feat(rulebooks): aethis rulebooks set-logic — set the composition expression on a rulebook. The composition expression (server field outcome_logic) is an Expr AST that combines per-ruleset outcomes into the rulebook’s final decision. Previously settable only via raw PATCH; now exposed via the CLI for multi-ruleset rulebooks (e.g. UK FSM’s child_eligibility AND (household_criteria OR universal_infant)).
    • aethis rulebooks set-logic <id> -f logic.yaml — load from YAML/JSON file
    • aethis rulebooks set-logic <id> --logic '<json>' — inline JSON
    • Exactly one of --file / --logic is required; both forms reject non-object payloads at the client side so server validation isn’t the first line of defence.
aethis-cli
2026-05-21
  • feat(rulesets): ruleset lifecycle commands scoped to a rulebook. Phase B.1b of the converged 2-term model. Adds four new sub-commands under aethis rulesets:
    • aethis rulesets list <rulebook> — list rulesets in a rulebook (grouped by ruleset_name with version counts, live version, and observed states). The legacy -p <project_id> and --public modes are preserved while the project-scoped authoring pipeline retires in a future phase.
    • aethis rulesets create <rulebook> <ruleset_name> [-n "Display name"] — create a new draft Ruleset inside the rulebook. The display name auto-derives from ruleset_name if not provided (child_eligibility → Child Eligibility).
    • aethis rulesets show <rulebook> <ruleset_name> — full version history for one ruleset name (bundle_id, version, state, created), with live version highlighted.
    • aethis rulesets promote-to-live <rulebook> <ruleset_name> <ruleset_id> [--note "..."] — atomically promote a testing-state ruleset version to live via the Phase A.4 service. Auto-cuts a new rulebook version; previous live ruleset is archived.
  • feat(client): four new AethisClient methods — create_ruleset_in_rulebook, list_rulesets_in_rulebook, show_ruleset_in_rulebook, promote_ruleset_to_live.
  • Requires aethis-core v0.20.0+ live on the target API (Phase A.8 endpoints).
aethis-cli
2026-05-21
  • feat(rulebooks): new aethis rulebooks command group. First user-facing surface for the converged 2-term authoring model (workspace PR #64, aethis-core PRs #133-139). A Rulebook is the whole form — the execution unit — that owns a locked field vocabulary, composition logic, rulebook-level test cases, and an integer version history.
    • aethis rulebooks list — list tenant rulebooks
    • aethis rulebooks show <id-or-slug> — full configuration
    • aethis rulebooks create <name> --domain <d> [--slug ...] — create draft
    • aethis rulebooks set-fields <id> -f fields.yaml — replace locked vocabulary
    • aethis rulebooks lock-fields <id> / unlock-fields <id> / get-fields <id>
    • aethis rulebooks tests add <id> -f scenario.yaml — embed full-form test case
    • aethis rulebooks tests list <id> / delete <id> <tc_id>
    • aethis rulebooks activate <id> / archive <id> — lifecycle
    • aethis rulebooks decide <id> -i '{...}' [--explain] — evaluate composed rulebook
    • aethis rulebooks schema <id> / explain <id> — combined schema + explanations
  • feat(client): new AethisClient methods for every rulebook REST endpoint (create / list / show / update / activate / archive / set-fields / lock-fields / unlock-fields / get-fields / add-test / list-tests / delete-test / decide-rulebook / get-rulebook-schema / explain-rulebook).
  • Requires aethis-core v0.19.0+ live on the target API (the Phase A.6 endpoints).
  • The legacy aethis projects / aethis generate / aethis test / aethis publish command tree is unchanged in this release — replacement lands in the next minor (Phase B.1b: ruleset lifecycle + project retirement). No backward-compat shims are planned past public release.
aethis-sdk-python
2026-05-20
  • feat(models): add name: Optional[str] to ruleset response models — surfaces the human-readable section name introduced in aethis-core v0.18.0. Adds RulesetSummary (anonymous catalogue / GET /api/v1/public/rulesets) and RulesetListItem (project-scoped / GET /api/v1/public/projects/{id}/rulesets) as typed models, and adds the same name field to SchemaResponse. Backwards-compatible: pre-backfill rulesets serialise with name=None.
aethis-mcp
2026-05-20
Add optional name parameter to aethis_publish tool — lets clients override the human-readable section name when publishing a ruleset. Companion to aethis-core v0.18.0’s PublishRequest.name field.

Added

  • aethis_publish tool now accepts an optional name parameter. When supplied, it overrides the section name stored on the ruleset (default is a titlecase of section_id, e.g. "english_language" → "English Language"). Section names are surfaced in rulebook responses so end users can see which sections compose a rulebook.
  • AethisClient.publish() now accepts a third name?: string argument and includes it in the POST body when set.

Tests

  • server.test.ts — two new aethis_publish cases: forwarding name to the client and echoing it in output; confirming name is omitted when not provided.
  • client.test.ts — two new publish() cases: body contains name when provided; body is absent when neither label nor name is set.
aethis-mcp
2026-05-20
Surface the human-readable section name in aethis_list_rulesets and aethis_discover_rulesets tool output. The engine has been returning name on both RulesetSummary (public catalogue) and RulesetListItem (project-scoped) responses since aethis-core v0.18.0; the MCP server already forwarded every API field verbatim via JSON.stringify, so the data was reaching the LLM, but the tool descriptions didn’t advertise the field. The descriptions now mention name so models know to read and surface it to users (e.g. “Knowledge of language and life in the UK” instead of just b_123…).

Changed

  • aethis_list_rulesets tool description now mentions the human-readable name field returned alongside ruleset ID, status, version, field count, and rule count.
  • aethis_discover_rulesets tool description now lists name in the documented response shape.

Tests

  • aethis_list_rulesets and aethis_discover_rulesets server tests assert the name field passes through to the LLM-facing JSON output.
aethis-cli
2026-05-20
  • feat(rulesets): show the human-readable section name column in aethis rulesets list output (both the public showcase and project-scoped tables). Surfaces the new field from aethis-core v0.18.0.
aethis-cli
2026-05-20
  • feat: pluggable auth providers. Profiles now carry an optional auth_mode (default "api_key") and audience field. The new aethis_cli.auth_providers module exposes a process-local registry; plugins (e.g. aethis-cli-internal) can register_provider("gcloud_id_token", ...) to add staff/internal auth schemes without touching the published package. AethisClient accepts an optional auth_provider callable, and make_authed_client(...) picks the right provider based on the active profile’s mode.
  • feat: aethis status now prints the active profile name + auth mode (plus audience when set). For non-api_key modes it shows “provider-minted at request time” instead of calling /me, which is X-API-Key-only.
  • chore: un-hide the --base-url global flag in aethis --help (it was already implemented, just hidden=True).
aethis-mcp
2026-05-19
Security hardening pass. Bundles the v0.5 security review fixes into one release. Closes #33, #34, #35; addresses GHSA-ph7q-r9q4-922g (disclosed on publish).

Security

  • GHSA-ph7q-r9q4-922g (high) — prompt injection via unsanitised API response text in aethis_explain_failure. formatExplainFailure now wraps every API-supplied free-text field (diagnosis, dsl_hint, criterion title / rule_text / source_refs) in an <api_response> fence and prepends a one-line preface telling the model the contents are data, not instructions. Literal closing tags inside payloads are neutralised so a payload cannot break out of the fence. The same fenceUntrusted helper has been applied to other free-text API surfaces (aethis_next_question, aethis_discover_sections, aethis_refine_sections, aethis_discover_fields, formatTestResults).
  • #33 — src/credentials.ts now resolves the credentials file via fs.realpath and asserts the canonical path sits under $HOME (or under an absolute XDG_CONFIG_HOME the user controls); refuses with UnsafeCredentialsError otherwise. Permissions check now matches ssh / aws-cli behaviour: any group/other bit set on the credentials file → refuse with Permissions 0NNN ... too open. Run: chmod 600 <path>.
  • #34 — progress_detail from the polling API is sanitised before it lands on stderr: control characters are stripped (TAB preserved) and the body is capped at 120 visible chars + …. Full-fidelity output is gated behind AETHIS_MCP_VERBOSE=1. Prevents the server from injecting terminal escape sequences or PII into anything that captures the MCP process stderr.

Changed

  • #35 — Authoring tools now accept safer per-call key forms.
    • New anthropic_key_env: string — name of an env var the MCP server reads at call time. Preferred. The raw value never appears in the tool call so it does not land in the MCP host’s session transcript JSONL.
    • New anthropic_key_keychain: string — macOS keychain reference, either service:account or just account (service defaults to aethis-anthropic-key).
    • Raw anthropic_key / openai_key arguments remain accepted for backwards compatibility but are now marked [sensitive — do not echo or log] in the schema; deprecated in tool descriptions.
    • resolveLlmKey (exported from src/credentials.ts) consolidates the resolution chain and throws MissingLlmKeyError if every form is empty.
    • Applies to aethis_generate_and_test, aethis_refine, aethis_discover_fields, aethis_refine_fields, aethis_discover_sections, aethis_refine_sections.

Docs

  • README: new “Passing your Anthropic key safely” section showing env / keychain forms first; raw-key form marked deprecated.
  • CLAUDE.md: new gotchas covering safe-key resolution and the untrusted-content fencing helper.
aethis-cli
2026-05-19
  • fix(decide): aethis decide --explain no longer crashes with AttributeError: 'str' object has no attribute 'get'. The CLI previously treated the engine’s explanation field as a flat list[dict], but the public decide route returns a layered {decision, decision_path?, groups: [{group, status, criteria: [{title, status, supporting_facts?, ...}]}], unused_facts} shape. The “Rules” block now walks the actual structure and renders each group + criterion with PASS/FAIL marks, supporting fact field/value pairs underneath satisfied criteria, and a final list of unused fields (provided answers that no satisfied criterion referenced — useful for catching field-name typos).
aethis-cli
2026-05-19
  • fix(login): default AETHIS_CLERK_CLIENT_ID to the OAuth Application registered on the clerk.aethis.ai Clerk instance. The previous default belonged to a different Clerk app, so aethis login returned invalid_client against the dev-tools domain set in 0.12.1.
  • fix(account): default AETHIS_CLERK_DOMAIN to clerk.aethis.ai for aethis account generate (matching the 0.12.1 change to aethis login); previously still pointed at the immigration domain.
aethis-cli
2026-05-13
  • feat: decide, explain, and fields no longer prompt for sign-in when no API key is present. Public rulesets are now accessible with zero setup — the CLI silently uses an anonymous client and lets the server return an error only if a private ruleset is requested.
  • fix: hide --base-url global flag from aethis --help (internal dev override; AETHIS_BASE_URL env var unchanged)
  • docs: reorder aethis --help to lead with the no-auth explore flow, then authoring
aethis-mcp
2026-05-12
  • docs: surface aethis-skills as the optional agent workflow layer on top of MCP.
aethis-cli
2026-05-12
  • fix: default Clerk domain changed from clerk.aethis.legal to clerk.aethis.ai so developer portal users can authenticate via aethis login (closes aethis-cli#40)
aethis-mcp
2026-05-11
  • fix: align package.json repository metadata with GitHub provenance so npm Trusted Publishing can verify the package source.
aethis-mcp
2026-05-11
  • fix: pin zod to v3 so the MCP SDK tool registration types match the build-time schema shape; npm publish now runs the prepublishOnly TypeScript build successfully.
aethis-sdk-python
2026-05-10
  • fix: update examples/session.py to use AETHIS_RULESET_ID env var (was deprecated AETHIS_BUNDLE_ID) and replace stale internal default with the public aethis/construction-all-risks slug
aethis-mcp
2026-05-10
  • docs: fix stale bundle/bundle_id/aethis_create_bundle terminology in docs/demo-construction-insurance.md and docs/agentic-decision-systems.md — these files were not caught by the v0.3.0 rename sweep. All references now use ruleset/ruleset_id/aethis_create_ruleset
  • chore: bump server.json version to 0.4.1 (was lagging behind package.json)
  • security: regenerate package-lock.json — bumps hono 4.12.12 → 4.12.18, fast-uri 3.1.0 → 3.1.2, ip-address 10.1.0 → 10.2.0, postcss 8.5.8 → 8.5.14; clears all 6 open Dependabot alerts (closes #19)
aethis-mcp
2026-05-10
  • feat: new aethis_discover_rulesets tool — lists the cross-tenant public showcase catalogue (no authentication required). Mirrors the no-auth policy of aethis_decide / aethis_schema / aethis_explain. Returns slug, ruleset_id, description, field_count, rule_count for each entry; the slug or ruleset_id can then be passed to the existing decision tools. Distinct from aethis_list_rulesets, which remains tenant-scoped and authenticated. Tool count: 24 → 25.
  • feat: client.discoverRulesets(limit, offset) wrapping GET /api/v1/public/rulesets.
  • docs: aethis-decide prompt and server-instructions now point at aethis_discover_rulesets for first-time discovery (no key) before falling back to aethis_list_projects → aethis_list_rulesets for authenticated tenant browsing.
aethis-cli
2026-05-10
  • fix: remove examples/demo_core.sh (internal dev script referencing aethis-core by name and a private API path — not intended for public release)
  • fix: update tests/e2e/test_spacecraft_e2e.py to resolve the spacecraft fixture from examples/spacecraft-crew-rules/ instead of an internal path; drop internal service name from comment
  • docs: fix “rule bundle” → “ruleset” in examples/spacecraft-crew-rules/README.md
aethis-cli
2026-05-10
  • feat(updater): gh-style update-check banner. On startup the CLI kicks off a background thread that queries PyPI; if a newer release is available it prints a one-line notice to stderr at exit: “A new release of aethis-cli is available: 0.11.0 → 0.12.0 — to upgrade, run: <method-aware command>”. Detects whether the install came via uv tool, pipx, or pip and renders the matching upgrade command. Result is cached for 24 h at ~/.config/aethis/update_check.json. Suppressed automatically when stderr is not a TTY (CI, piped output). Disable with AETHIS_DISABLE_UPDATE_CHECK=1. The check never blocks the command — failures are silent.
aethis-cli
2026-05-10
  • feat(rulesets): aethis rulesets list --public lists the cross-tenant public showcase catalogue (no auth required). When run with no --project-id and no project context, falls through to the public catalogue automatically with a one-line hint — so a fresh signup sees something the moment they install the CLI instead of an empty list. Combine with aethis fields -b <slug> / aethis explain -b <slug> / aethis decide -b <slug> to fully exercise a ruleset without an API key.
  • feat(profiles): named credential profiles with both per-invocation flag (aethis --profile new-dev …) and sticky default (aethis profile use new-dev). Manage with aethis profile list/use/add/remove. Reserved profile name anonymous forces unsigned mode — handy for testing what a fresh signup sees without losing your admin key. aethis login --profile <name> writes into the named slot. Credentials file format upgraded to {active_profile, profiles: {...}}; legacy single-key files are read transparently and rewritten to the new shape on next save.
  • feat(client): AethisClient(unsigned=True) and make_anonymous_client() helper for paths that must hit the anonymous surface without accidentally sending a cached key.
  • feat(client): client.list_public_rulesets(limit, offset) wrapping GET /api/v1/public/rulesets.
aethis-cli
2026-05-08
  • feat(publish): thread --force through to the server-side TDD gate introduced in aethis-core 0.11.0. client.publish() gains a force_unsafe: bool = False keyword; aethis publish --force now passes force_unsafe: true in the request body so the server-side gate is bypassed (and a publish_force_bypass audit event is recorded). Older engines ignore the field — no breakage. Without --force, the new gate refuses publishing over a failing test suite even when the CLI’s own test gate is bypassed (e.g. by a direct curl that doesn’t use the CLI). Closes the cli/server asymmetry that nearly shipped a 10/11 ruleset to a canonical aethis/* slug on 2026-05-07.
aethis-sdk-python
2026-05-07
  • docs: link to the test-driven authoring guide on docs.aethis.ai and surface the publish-gate guarantee (rulesets cannot publish with a failing test) in the private-beta callout. Reference surface only — no code changes
aethis-mcp
2026-05-07
  • docs: surface the test-gate guarantee — aethis_publish refuses to publish a ruleset with a failing test, derived from positioning bible §5/§7. Strengthens the existing Note to an Important callout and annotates the publish line in the four-stage workflow
  • docs: drop force=true mention from troubleshooting — surfacing the override on the public README undermines the “cannot be published with failing tests” guarantee. The API parameter remains in the engine; whether to deprecate it is tracked separately
  • docs: fix tool count (25 → 24); tools table sums to 24 (5 + 7 + 8 + 2 + 2). Fixed in README header and in CLAUDE.md
aethis-mcp
2026-05-07
  • docs: link to docs.aethis.ai/agents/onboarding from Install section
aethis-cli
2026-05-07
  • docs: link to docs.aethis.ai/agents/onboarding from MCP one-liner section
aethis-sdk-python
2026-05-06
  • docs: remove positioning paragraph above Install — reference surface (per aethis.os/positioning/surface-types.md); the tagline is enough
aethis-sdk-python
2026-05-06
  • docs: add private-beta callout for authoring endpoints (decision endpoints remain anonymous)
aethis-sdk-python
2026-05-06

Changed

  • docs: align README with positioning bible — add problem/solution/methodology intro paragraph before Install section.
  • docs: add aethis-bible: markers to derived copy blocks (sourced from public-messaging.md §3/§4).
  • fix: terminology audit found no deprecated “rule bundle” or <5ms instances in README; no replacements needed.
aethis-sdk-python
2026-05-06

Changed

  • Aethis(api_key=...) and AsyncAethis(api_key=...) now accept api_key=None (or no argument) for the developer beta. Evaluation endpoints (/decide, /schema, /explain, /source) work anonymously, so the SDK no longer forces a key on instantiation. When api_key is omitted, the x-api-key header is simply not sent. Authoring endpoints will still return 401 without a key. Existing callers passing api_key="..." are unaffected.
  • README quickstart now shows Aethis() (no key) as the primary form, targets aethis/uk-fsm/child-eligibility (a live public ruleset) instead of the dated eng_lang:20250912-ec5d7c23, and prints the audit fields (inputs_hash, decision_id, decision_time, engine_version) added in 0.3.2. Configuration table updated: api_key is now documented as optional during the developer beta.
  • examples/oneshot.py refreshed to match: no key required by default, AETHIS_BUNDLE_ID env var renamed to AETHIS_RULESET_ID (catching the 0.3.0 bundle → ruleset rename it had missed), targets the live UK Free School Meals ruleset, prints the audit fields.

Notes

  • Backwards-compatible: Aethis(api_key="ak_live_...") continues to work exactly as before.
  • This pairs with the public-surface positioning that evaluation is free during the developer beta — see docs.aethis.ai.
aethis-sdk-python
2026-05-06

Added

  • DecideResponse.decision_id — per-call audit identifier returned by the engine.
  • DecideResponse.inputs_hash — canonical SHA-256 fingerprint of the input set.
  • DecideResponse.decision_time — ISO-8601 timestamp of the decision.
  • DecideResponse.engine_version — aethis-core@<semver> string identifying the engine that produced the decision.

Fixed

  • DecideResponse previously declared ruleset_id twice; Pydantic silently overrode the first declaration with the second. Deduplicated.
  • The four audit fields above were already returned by /api/v1/public/decide but were silently dropped by Pydantic because the model didn’t declare them. Callers can now read them directly off the typed response — no need to reach for the raw JSON. This is the audit-trail fingerprint that the docs and homepage prominently advertise (inputs_hash, decision_id); shipping an SDK that hid it was a defect.

Notes

  • Backwards-compatible. All four new fields default to None, so older engines that don’t emit them still parse cleanly.
aethis-sdk-python
2026-05-06

Fixed

  • aethis_sdk.__version__ now resolves from installed package metadata via importlib.metadata instead of a hardcoded constant. Previously reported "0.1.0" on every install regardless of the actual package version. Falls back to "0.0.0+unknown" only if the package is imported without being installed (editable dev or zip-on-PYTHONPATH).
  • Package description on PyPI: "…and bundle schemas" → "…and ruleset schemas" to match the v0.3.0 public-surface rename.

Added

  • README PyPI / Python-version / License shields.
aethis-mcp
2026-05-06
  • docs: restructure README as dev MCP docs — Install / Quick start / Tools / Setup leads, positioning sections (Problem, Accuracy, When to use this, How it works, Example walkthrough) removed; their content belongs in docs.aethis.ai or the benchmarks repo
  • docs: trim narrative paragraphs across Quick start, Conversational eligibility, and Authoring; collapse repeated Tips into terse callouts
  • docs: header tagline rewritten to a single factual line; link bar updated to the new structure
aethis-mcp
2026-05-06
  • docs: normalise tone to documentation register — replace argumentative Proof section with one-liner accuracy claim, neutralise example framing, trim sales-y bullets in When to use this
  • docs: add private-beta callout for authoring tools (decision tools remain public, no key required)
aethis-mcp
2026-05-06
  • docs: align README with positioning bible — promote 225-scenario accuracy framing
  • docs: add aethis-bible: markers to derived copy blocks
  • docs: fix latency claim to <1ms (was <5ms)
  • fix: replace deprecated “rule bundle” terminology with “ruleset”
aethis-cli
2026-05-06
  • docs: remove Why Aethis section — package README is a reference surface (per aethis.os/positioning/surface-types.md); install / quick start / authentication is the right lead, not a problem statement
aethis-cli
2026-05-06
  • docs: add private-beta callout for authoring tools (decision tools remain public, no key required)
  • docs: clarify in Authentication that aethis login requires an invite during the beta
aethis-cli
2026-05-06
  • docs: align README with positioning bible — add Why Aethis section, solution framing, TDD methodology beat
  • docs: add aethis-bible: markers to derived copy blocks
  • fix: replace deprecated “rule bundle” terminology with “ruleset” in pyproject.toml description
aethis-sdk-python
2026-05-05

Changed (Breaking)

  • Renamed the public bundle concept to ruleset throughout the SDK to match the aethis-core 0.10.0 API contract. Every bundle_id parameter and JSON key is now ruleset_id. URL paths inside the client moved from /api/v1/public/bundles/... to /api/v1/public/rulesets/.... The Session constructor now takes ruleset_id and exposes session.ruleset_id instead of session.bundle_id. Class names: BundleSummary → RulesetSummary.

Required

  • Engine aethis-core 0.10.0 or newer. Older engines respond at the legacy /bundles/* paths and this client will 404. Pin aethis-sdk==0.2.0 to keep working against an older engine.
aethis-mcp
2026-05-05
  • Breaking: renamed the public bundle concept to ruleset throughout the MCP tool set, to match the aethis-core 0.10.0 API contract. The compiled rule artefact is now called a ruleset in every tool name, parameter, and prose description. Specifically:
    • Tools: aethis_create_bundle → aethis_create_ruleset, aethis_list_bundles → aethis_list_rulesets, aethis_archive_bundle → aethis_archive_ruleset
    • Parameters: every bundle_id → ruleset_id
    • JSON keys returned to the agent: bundle_id/latest_bundle_id/bundle_version/deprecated_bundles/result_bundle_id/bundle_refs → ruleset_id etc.
    • URL paths inside the client: /bundles/... → /rulesets/...
  • This release requires aethis-core 0.10.0 or newer. Older engines respond at the legacy /bundles/* paths with bundle_id JSON keys; this client expects /rulesets/* and will 404. Pin aethis-mcp@0.2.6 if you need to keep working against an older engine until you can deploy.
  • MCP tool renames are part of the public LLM-facing contract. Coding agents that have learnt the old tool names (aethis_list_bundles etc.) from training data will get “no such tool” errors and need to retry against the new names. Tool descriptions explicitly call out the new naming so the LLM picks it up on first read.
aethis-cli
2026-05-05
  • Breaking: renamed the public bundle concept to ruleset throughout the CLI to match the aethis-core 0.10.0 API contract. The compiled rule artefact is now called a ruleset everywhere — in command names, in flag names, in JSON keys, and in prose. Specifically:
    • aethis bundles list/archive → aethis rulesets list/archive
    • --bundle-id flag → --ruleset-id
    • client.list_bundles() / archive_bundle() / get_bundle_schema() / explain_bundle() / get_bundle_source() / set_bundle_visibility() SDK methods → *_ruleset
    • JSON keys bundle_id / latest_bundle_id / bundle_version / bundle_refs → ruleset_id etc.
    • Default scope strings bundles:read/explain/write → rulesets:* (validated against the engine’s permission registry)
  • This release requires aethis-core 0.10.0 or newer. Older engines return bundles:* scopes and the CLI will reject them as invalid. Pin to aethis-cli==0.7.2 if you need to keep working against an older engine until you can deploy.
aethis-mcp
2026-05-03
  • Docs: replaced two stale aethis.ai/sign-up request-access pointers in the README authoring section with aethis.ai/developer-access. After the Clerk cutover, /sign-up serves the Clerk SignUp form for invitees rather than the Notion request-access form. No code or behaviour changes.
aethis-cli
2026-05-03
  • Docs: replaced the stale aethis.ai/sign-up request-access link with aethis.ai/developer-access in the README “Author your own rules” section and in the aethis whoami hint shown when the active key has no authoring scope. After the Clerk cutover, /sign-up serves the Clerk SignUp form for invitees rather than the Notion request-access form, so external “Request access” pointers were broken. No code path changes.
aethis-mcp
2026-05-01
  • Docs: README Quick start now leads with aethis mcp install --target all (via aethis-cli v0.5.0+). The manual claude mcp add and per-client JSON tabs are demoted to “Manual install” beneath. Setup section gains a Keys & security subsection covering AETHIS_API_KEY vs ANTHROPIC_API_KEY placement (MCP client config, not shell), rotation workflow (aethis account generate + aethis account revoke), and multi-machine guidance.
  • Discoverability: package.json keywords extended with regulation, policy, eligibility-check, deterministic-decision — matches the highest-intent search terms used by developers in regulated domains. Existing keywords retained.
  • CLAUDE.md updated to note the aethis mcp install install path so future contributors don’t re-document the manual JSON as primary.
No code or behaviour changes.
aethis-cli
2026-05-01
  • Docs: README gains a dedicated Authentication section explaining the three modes (aethis login for explicit setup, lazy auth for inline mid-command sign-in, --no-prompt for CI). Authoring quickstart leads with aethis init (the v0.7.0 wizard prompts for a name and runs sign-in itself, so aethis login as a separate step is no longer needed). Environment-variable table expanded to cover AETHIS_BASE_URL and ANTHROPIC_API_KEY. Troubleshooting entry for Auth error now mentions the lazy-auth prompt and --no-prompt. CLAUDE.md updated to document the aethis mcp install path, lazy-auth helper, and --no-prompt flag for future agents working on the CLI. No behaviour change.
aethis-cli
2026-05-01
  • New: aethis init first-run wizard. With no args, prompts for the project name (default = current directory name); a positional aethis init <name> keeps working unchanged. If no API key is cached, triggers the same OAuth flow as aethis login before any filesystem writes — Ctrl-C during browser sign-in no longer leaves a half-scaffolded project on disk. After scaffolding, prints the next-step ladder (aethis sections discover → fields discover → generate --poll) so new users have a clear path forward. New --no-prompt flag for scripted use; with that flag, missing required values fail fast and missing auth surfaces a clean AuthRequired error instead of opening a browser. 10 new tests covering prompted, non-prompted, no-auth + interactive, no-auth + --no-prompt, and name-validation paths. Closes #15.
aethis-cli
2026-05-01
  • New: lazy auth. Authenticated commands (aethis projects list, generate, publish, etc.) now detect missing credentials or 401 responses and offer an inline browser sign-in prompt: "No API key. Open browser to sign in? [Y/n]". On accept, the same OAuth flow as aethis login runs, the key is cached, and the original command retries — exactly once, no infinite loops. Non-TTY stdin/stdout (CI, pipes) and the new --no-prompt global flag skip the prompt and surface a clean AuthRequired error. --api-key <key> still bypasses the helper entirely. New helper module aethis_cli/auth_helpers.py; the OAuth flow inside commands/login_cmd.py was factored into a reusable run_browser_login(). 17 new tests in tests/test_lazy_auth.py. Closes #12.
aethis-cli
2026-05-01
  • New: aethis mcp install --target <client> writes the MCP server entry into your editor’s config in one shot. Supports claude-code (project-level .mcp.json), cursor (~/.cursor/mcp.json), claude-desktop (~/Library/Application Support/Claude/claude_desktop_config.json on macOS, ~/.config/Claude/... on Linux), windsurf (~/.codeium/windsurf/mcp_config.json), and --target all for everything at once. Idempotent, preserves any other configured MCP servers. aethis mcp uninstall --target <client> reverses the install. Closes #16.
aethis-cli
2026-05-01
  • UX: aethis login --help now reads “Sign in and store an API key locally. First-time setup — this is all you need.” aethis account generate --help clarifies it’s for additional keys (rotation, multi-machine, scoped access). After successful aethis login, a tip line points at aethis status / aethis account keys. README quickstart collapses any “first login then generate” sequence into a single aethis login step. No behaviour change. Closes #13.
aethis-cli
2026-05-01
  • Docs: README install section now leads with uv tool install aethis-cli (recommended) and pipx install aethis-cli, with pip install in a venv as the third option. Pairs with Aethis-ai/docs#12. Closes #14.
aethis-mcp
2026-04-28
First version published to npm since 0.2.2. The v0.2.3 tag exists in git but predates the publish workflow — it never reached npm. This release rulesets all work since 0.2.2.

Registry

  • MCP Registry submission ready. Added mcpName: io.github.aethis-ai/aethis-mcp to package.json and a top-level server.json declaring the npm package, transport, and environment variables. Submit via mcp-publisher after npm publish.

Breaking Changes

  • openai_key parameter renamed to anthropic_key on aethis_generate, aethis_generate_and_test, and aethis_refine. The old parameter name is still accepted for backwards compatibility but will be removed in a future release.

Improvements

  • Better error messages on generation failure. Failed jobs now surface classified error details (invalid key, rate limit, connection failure) instead of “unknown error”.
  • Sends both X-Anthropic-Key and X-OpenAI-Key headers for backwards compatibility with older API versions.
  • aethis_explain_failure clarification. Tool docs now note that ruleset_id must be the concrete ID from a /decide envelope; slugs are not yet resolved on this endpoint (tracked in aethis-core#51).

Docs

  • Proof section updated to cite the Simpson et al. 2026 benchmark paper. Replaced the pre-paper 11-scenario table (GPT-5.4-mini 82%, GPT-5.3 27%) with paper-backed figures from Table 8b of the published benchmark. Removed the 27% GPT-5.3 claim — the paper identifies that figure as a harness-configuration bug; the corrected value is 63.6%.
  • Proof section: add §6.10 LegalBench external-validation paragraph. v3.8 of the paper adds external validation across 9 LegalBench tasks (949 held-out cases). Combined paired-binomial McNemar’s: p < 0.001 vs Sonnet 4.6, p = 0.003 vs Opus 4.7, p < 0.001 vs GPT-5.4. Linked to the public LegalBench harness at confidently-wrong-benchmark/legalbench/.
  • Proof section: replaced 11-scenario subset table with v3.8 adversarial extension (§6.4.1). The v3.7 11-scenario exception-chain table no longer differentiates current frontier models from the engine (GPT-5.4 default and low both 11/11, Opus 4.7 11/11). The Proof section now leads with the v3.8 adversarial extension (20 newly-authored scenarios; engine 20/20; Opus 4.7 18/20; GPT-5.4 default 19/20 with 0 reasoning tokens; Sonnet 4.6 19/20) and the shifting-ground argument from paper §6.5 Finding 6.
  • Use aethis/construction-all-risks slug in CAR proof example for stable URL across ruleset regenerations.
  • Invite-only beta messaging replaces “rolling out now” framing throughout README — explicit approval-gated framing aligned with current onboarding.
  • docs.aethis.ai badge added to README.

Internal

  • Added .github/workflows/publish.yml (provenance via OIDC + NPM_TOKEN) so future tag pushes auto-publish.
  • Added Claude PR review workflow (dry-run mode).
  • Added internal CLAUDE.md for agent onboarding.
aethis-cli
2026-04-28
Two bug fixes that block the documented quickstart against public bundles.

Bug fixes

  • aethis decide -b <slug> / explain -b <slug> / bundles archive -b <slug> now accept slugs. The classifier in _id_utils.classify_id previously returned "unknown" for slugs (e.g. aethis/uk-fsm/universal-infant), and require_bundle_id rejected them with "is not a valid Bundle ID". The public API resolves both bundle IDs and slugs on /decide, /schema, and /explain, so the CLI now passes both through. Error message updated to mention slugs and link to aethis bundles list.
  • aethis fields -b <bundle> no longer requires an aethis.yaml. It now uses the same load_client_or_fallback() helper as decide, explain, bundles, and projects — read-only commands work from any directory. Previously this command errored out with "No aethis.yaml found" even when called with a concrete bundle reference.
aethis-sdk-python
2026-04-27

Added

  • DecideResponse.slug — stable, human-readable handle for the ruleset (e.g. aethis/uk-fsm/child-eligibility). Set when the resolved ruleset was published under a slug; None otherwise. Prefer this over ruleset_id for any reference that should survive ruleset regeneration.
  • SchemaResponse.slug — same handle, surfaced from GET /rulesets/{id}/schema.

Notes

  • Backwards-compatible. Existing code that reads ruleset_id keeps working unchanged; slug is purely additive.
  • Requires the aethis-core engine release that surfaces the field in /decide and /rulesets/{id}/schema responses (rolling out 2026-04). Older engines will simply leave slug=None.
aethis-cli
2026-04-19

aethis status output polish

  • Server line now shows just the URL when it’s the default (https://api.aethis.ai) — the (default — no override) suffix was noise in the common case. Overrides (AETHIS_BASE_URL, aethis.yaml) still show source with a green marker.
  • Identity line now says ✗ API key rejected (run \aethis login` to re-authenticate)when/mereturns 401/403/404, instead of the raw✗ 404 from /me (Not Found)` HTTP message. Other HTTP errors keep a contextual message.
aethis-cli
2026-04-19
This release ships the rich-status and read-only-from-anywhere work that the 0.2.0 notes already described but which hadn’t actually been merged into a published release yet. (The code was sitting in a local branch; the prior 0.2.x/0.3.x wheels still had the minimal status command.)

aethis status — context-aware summary

  • aethis status with no args now prints CLI version, resolved server URL (with source — env / yaml / default), loaded aethis.yaml, bundle id from .aethis/state.json, and whoami identity (key id, tenant, tier, scopes, can_author). Helps answer “what will my next command actually hit?” before running it.
  • aethis status -p <project_id> (or from inside a project dir) still shows generation progress, appended after the global summary.

Read-only commands usable from anywhere

  • aethis explain, decide, bundles list, bundles archive, projects list, projects show, projects archive no longer require an aethis.yaml in the current directory — they fall back to AETHIS_BASE_URL (or the default https://api.aethis.ai).
  • aethis explain / decide now reject Project IDs (proj_*) passed to -b/--bundle-id with a one-line hint pointing at the Bundle column of aethis projects list, instead of silently 404’ing.

Internals

  • New resolve_base_url_with_source() / load_client_or_fallback() helpers in aethis_cli/config.py that the above commands share.
  • New aethis_cli/commands/_id_utils.py + test coverage for bundle-id validation.
  • New tests for explain, status, and _id_utils.
aethis-cli
2026-04-19

Docs cleanup

  • README and docs.aethis.ai/interfaces/cli no longer document AETHIS_BASE_URL or show base_url: in the aethis.yaml example — public users always hit https://api.aethis.ai, and the documented values were just duplicating the default. The env var still works as an override for devs and CI; it’s intentionally undocumented.
  • Dropped the AETHIS_CLERK_DOMAIN env var from the README (marked “development only” and confusing for public users). The override still works in code.
aethis-cli
2026-04-19

Trim public CLI to the developer API surface

The public CLI now only ships commands every developer can use against https://api.aethis.ai. Privileged and staff-only commands have been removed and will live in a separate internal plugin package.Breaking changes:
  • Removed aethis source — internal-only DSL viewer; moved to the aethis-cli-internal plugin.
  • Removed aethis account permissions — IAM permission registry; internal-only.
  • Removed the aethis guidance domain … group (and the deprecated aethis domain guidance … alias) — domain-level guidance is staff-managed.
  • Removed the global --base-url flag (plus the per-command --base-url on login, account generate, account keys, account revoke). The AETHIS_BASE_URL env var still overrides the default. The flag had no meaning for the public API target and cluttered --help.
New: third-party plugin support.
  • The CLI now discovers plugins via Python entry points under the aethis_cli.plugins group. A plugin exposes one callable register(app: typer.Typer) -> None and attaches extra commands to the root app. Plugin load failures print a single warning to stderr and never crash the CLI.
  • The staff-facing aethis-cli-internal package uses this hook to re-attach source, domain guidance, permissions, and the --base-url flag.
aethis-cli
2026-04-19

Consolidated guidance command tree

  • aethis domain guidance ... moved under aethis guidance domain ... — the domain group exists only to host guidance, so having two top-level trees for the same concept was confusing. All four subcommands (add, list, import, export) behave identically on the new path.
  • The old aethis domain guidance ... path still works as a hidden deprecated alias: invocations continue to succeed and emit a one-line deprecation notice to stderr. It is no longer shown in aethis --help. Planned removal in a future release.
aethis-cli
2026-04-19

aethis status — global CLI context

  • New behaviour: aethis status (no args) now prints a one-screen summary of the current CLI context: CLI version, resolved server URL (with source — --base-url / env / yaml / default), loaded aethis.yaml + project, bundle id from .aethis/state.json, and whoami identity (key id, tenant, tier, scopes, can_author). Answers “what will the next command hit?” — the usual cause of “why is my project missing?” is talking to the wrong server.
  • Backward compatible: aethis status -p <project_id> (or invoked from a project dir) still shows generation progress, now appended after the global summary.

UX improvements for read-only commands

  • aethis explain, decide, bundles list, bundles archive, projects list, projects show, and projects archive no longer require an aethis.yaml in the current directory — they fall back to AETHIS_BASE_URL (or the default https://api.aethis.ai) when invoked from anywhere.
  • aethis explain and decide now reject Project IDs (proj_*) passed to -b/--bundle-id with a one-line hint pointing at the Bundle column of aethis projects list, instead of silently proceeding to a 404.
  • aethis --base-url <url> is now a top-level flag, equivalent to setting AETHIS_BASE_URL for one invocation. Lets you hit staging or a self-hosted instance without editing aethis.yaml.
  • aethis projects list prints a short tip after the table showing how to copy a Bundle value into aethis explain -b ….
  • Configuration and authentication errors now render as a single red line via the existing cli() handler, not a Rich traceback panel. pretty_exceptions_enable=False is set on every Typer app.

Better --help

  • Top-level aethis --help now shows common flows (status, list, explain, decide), authoring flow, and how to target a different server.
  • explain, decide, bundles list, projects list, and status all have “Examples:” blocks in their per-command help.
aethis-mcp
2026-04-14

New Tools

  • aethis_add_domain_guidance — Add cross-section guidance hints at domain level (e.g. uk_citizenship). Applies automatically to all projects in the domain during generation.
  • aethis_list_domain_guidance — List all active domain-level guidance hints.
  • aethis_list_guidance — List all guidance hints accumulated for a project. Use before adding new guidance to avoid duplicates.
  • aethis_explain_failure — Diagnose a failing test case. Returns criterion statuses with DSL metadata and a targeted fix hint.

Improvements

  • aethis_add_guidance now accepts process_type ("rule_generation" | "field_extraction"). Use field_extraction for field design principles (solicitor navigation, raw-facts principle). Defaults to "rule_generation".
  • aethis_add_domain_guidance accepts notes — SME commentary or legislation provenance stored on the hint. Never sent to the LLM.
  • Two-level hint retrieval: generation now fetches domain-level hints (cross-section) alongside project-level hints in a single pass.
aethis-mcp
2026-04-09

New Tools

  • aethis_discover_fields — Discover input fields from source text. Returns field names, types, and completeness assessment. Call before writing test cases.
  • aethis_refine_fields — Iterate on field discovery with targeted feedback.

Improvements

  • Added aethis-author and aethis-decide MCP prompts for compatible clients (Claude Desktop, Cursor, VS Code Copilot).
aethis-mcp
2026-04-05
Initial release.

Features

  • Decision tools: aethis_schema, aethis_decide, aethis_next_question, aethis_explain
  • Discovery tools: aethis_list_projects
  • Authoring tools (TDD workflow): aethis_create_ruleset, aethis_generate_and_test, aethis_add_guidance, aethis_refine, aethis_publish, aethis_archive_project, aethis_archive_ruleset
  • HTTPS enforcement for remote hosts
  • Exponential backoff with retry on 429/502/503/504
  • Works with Claude Desktop, Claude Code, Cursor, and Windsurf
Note: v0.1.0 used aethis_create_ruleset (renamed to aethis_create_ruleset) and aethis_project_status (replaced by aethis_list_projects). These tools were removed in v0.2.x.
aethis-cli
2026-04-05
Initial release.

Features

  • Account management: aethis account generate (browser OAuth), aethis account keys, aethis account revoke
  • Project authoring: aethis init, aethis generate --poll, aethis test, aethis publish
  • Decision tools: aethis decide, aethis fields, aethis explain
  • Project management: aethis projects list, aethis bundles list, aethis bundles archive
  • Security: HTTPS enforcement, OS keychain storage, PKCE OAuth flow
  • Example: Spacecraft Crew Certification Act 2049 with 5 golden test cases