CHANGELOG.md of the package it belongs to and is labelled with that package. Subscribe via the RSS feed for this page.
- ci: the post-publish unstick step queries every downstream repository at its current path. Two entries still named a pre-transfer path. They worked through redirects, but a new repository created at the old path would have been queried instead and returned no pull requests. No change to the CLI itself.
- ci: the post-publish unstick step now sees repositories that have been
renamed or transferred. It listed each downstream repository’s open pull
requests with a search filter, and the search index does not follow a
repository rename or transfer: a moved repository returned no results and a
success exit code, so it was skipped without a warning. The step now lists
open pull requests without the search filter and matches the
aethis-needsmarker in each body, which it already did. A repository that cannot be listed now produces a workflow warning, and so does one that hits the listing limit of 100 open pull requests. No change to the CLI itself.
Authoring safeguards are now visible in tool output. Both are warn-only and never block a generation or a publish.
aethis_generate_and_testandaethis_refinerender the source questions authoring raised: places where the source text conflicts with itself or can be read more than one way. Each question shows the quoted clauses with their citation keys, the candidate readings, and the provisional reading the ruleset encodes.aethis_publishrenders them too.aethis_publishrenders the source check: a warning when a cited document is not the text the ruleset was built from (mismatch), when a digest is unavailable (unverifiable), or when the ruleset predates input recording (no_authoring_inputs_recorded).- Fix: the generate-and-test path kept only the ruleset id and test result from the final generation status, so fields carried there (source questions and the run’s question counts) never reached the tool output. They are now preserved.
- All question and check text comes from uploaded sources and model output, so it is returned inside the
<api_response>untrusted-content fence. - Security hardening: the untrusted-content fence is harder to break out of. It previously neutralised only the exact closing tag
</api_response>. Now any occurrence of the fence name in returned text is neutralised, whatever surrounds it, so no spacing, case, lookalike bracket or slash, encoding, or forged opener such as<api_response label="system">can close the fence or open a new one. Fence labels are restricted to identifier characters, so returned data can no longer reach a fence label’s attribute unescaped. aethis_next_questionno longer prints a note’smetadata.typeas bare[type]text before the note. That value is returned data, so it now appears only in the note’s fence label.
- feat: print the source check and source questions.
aethis publish,aethis rulesets promote-to-live,aethis generate --pollandaethis statusnow print two warn-only fields when the engine returns them.source_checkcompares the bytes each citation resolves to at publish with the bytes the ruleset was generated from. The CLI prints it when the status iswarningsorerror: eachmismatchwith both digests, eachunverifiablecitation, andno_authoring_inputs_recorded. Anokornot_runcheck prints nothing.source_questionslists conflicting, ambiguous or missing source text that authoring raised rather than resolving silently. The CLI prints a count and, for each question, its kind, the quoted clauses, the provisional reading the ruleset uses, the affected criteria and, for a refine, the ruleset it was inherited from.- Responses without these fields print exactly what they printed before.
- Question and warning text comes from uploaded sources and model output. It
is never interpreted as terminal markup, and terminal control characters in
it (escape sequences, C1 controls, bidirectional overrides, Unicode line
separators, lone surrogates) are shown as
visible escapes such as
\x1b; embedded line breaks become spaces.
Security: provider credentials leave the process only when the user configured them for Aethis. Upgrade recommended.
- Breaking: an Anthropic key is read from the environment only via the new
AETHIS_ANTHROPIC_KEY_ENVserver setting, in which the user names the variable holding the key. Ananthropic_key_envvalue supplied in a tool call is refused unless it equals that configured name, so a host model can no longer opt a user’sANTHROPIC_API_KEY(or any other variable) in on its own. The missing-key error now tells the user how to configure a key instead of suggesting an argument for the model to retry with. - Breaking: the retired
openai_keyargument is refused. Previously it was sent in the Anthropic key header. - Any key value that is not Anthropic-shaped (
sk-ant-…) is refused locally and never sent. - Key-shaped text (
sk-ant-…,sk-proj-…, provider-masked echoes) in any upstream response or error is masked before it reaches the MCP client. - If you previously exported a provider key in your MCP host’s environment and used the authoring tools, rotate that key.
- chore(examples): retire the bundled spacecraft example in favour of the
maintained one.
examples/spacecraft-crew-rules/is removed: its copy of the Spacecraft Crew Certification Act had drifted from the canonical text. The README now points at the maintained example in Aethis-ai/aethis-examples. - test(e2e): the spacecraft authoring e2e really runs. It previously pointed
at a file path that did not exist and skipped silently. It now fetches the
Act, the scenarios and the guidance hints from the maintained example at a
pinned commit, verifies the Act’s digest, and fails (never skips) when a
fetch fails or the digest does not match. Its own hard-coded guidance and
test cases are removed; every maintained scenario is checked via
decide. Pinned to aethis-examples84cad29(v0.2.7), whose examplesources/directory holds only the canonical Act and its citation manifest. - ci(authoring-e2e-weekly): the lane revokes its key after a real run. The revoke step reused the mint step’s Clerk session token, which expires in about a minute. That went unnoticed only because the test used to skip instantly. It now signs in afresh before sweeping the lane’s keys. The lane also refuses to run pytest without a minted key, so a missing key can no longer turn the authoring tests into skips and the lane green.
- test: the pinned-example check runs on every PR. A credential-free test fetches the pinned example and verifies the Act’s digest in normal CI, so a broken pin is caught per-PR rather than weekly. The redundant 80% pass-rate assertion is removed; the strict per-scenario check covers it.
- fix(generate): preserve structured authored field notes on generation
pins.
fields.yamlnotes, including their opaque metadata, now reach the existing engine note contract. The CLI checks advertised engine support before upload and rejects known unsupported engines.
- Add
aethis_set_tests(project_id, test_cases)for destructive replacement of one existing project’s complete reviewed 1–100-case suite. It verifies the target OpenAPI replacement capability before writing, preserves project sources, fields and guidance, and never creates another project. - Replacement POSTs are sent once. Interrupted responses report an unknown outcome and require inspection before another approved replacement.
- Preserve immutable publication receipt metadata (
published_version_id,content_digest, andpublished_version_label) inaethis_publishoutput.
- Resolve the selected CLI profile’s API key and endpoint together, honoring non-secret
AETHIS_PROFILEand absoluteXDG_CONFIG_HOMEreferences. Profile credentials outrank stale legacy keychain entries; flat legacy files remain supported. - Preserve the default endpoint for standalone environment keys, including when the implicitly active CLI profile is anonymous. Explicit named-profile overrides retain the selected endpoint.
- Refuse invalid selected profiles and unsafe files visibly without logging secrets. Preserve explicit environment overrides and unsigned anonymous setup.
- Keep startup identity paired until restart; late login refuses an endpoint change before sending a request.
- feat(mcp): secure selected-profile setup for Codex and existing hosts.
aethis mcp installnow supportscodex(and includes it inall) through Codex’s nativemcp add/get/removecommands. Every host registration stores onlyAETHIS_PROFILEand an absoluteXDG_CONFIG_HOME, never an Aethis API key or endpoint. A clean install uses the unsignedanonymousprofile; saved profiles stay pinned even if the CLI active profile changes later. Conflicting one-off key or endpoint overrides and ambiguous user-managed registrations fail before host configuration is changed. Known legacy CLI registrations migrate safely.
Added
-
Immutable rulebook releases with two-word labels (#39, workspace
epic #1299 P1). Publishing a rulebook now freezes its complete runtime
content — every member resolved to exact compiled bytes (plus the published
leaf
version_id/content_digest), the composition, the field vocabulary overlay, conversational guidance (robot_hints), the scope declaration and the published title — into one immutableRulebookVersionrelease with its own UUID (release_id), a permanent two-word label (e.g.Amber Heron, allocated from small neutral word lists under a database uniqueness constraint) and acontent_identitydigest. An identical publication reuses its existing release and label; any change to published content mints a new one. The family’s “latest” is a CAS-protected head on the existing publication-head collection, so a publication is observed complete-old or complete-new, never mixed. -
The existing operations publish; there is no new endpoint:
promote-to-live(always),POST /rulebooks/{id}/activate(always, idempotent — also how a family transferred onto an engine without its releases is published there), andPATCH /rulebooks/{id}/POST /rulebooks/{id}/fieldswhen they supply published content on a published family. Draft families keep pure draft editing. A publication failure is a structured 422 (release_member_unresolved,release_label_allocation_failed, …). Member validation fails before the draft save; a later publication failure keeps the valid draft so the identical content request can resume the idempotent publication. -
Additive read surface:
POST /decideacceptsrelease_id(UUID or exact label,rulebook_idrequests only);GET /rulebooks/{id}/schema|explain|graphaccept?release_id=. Responses carry an additivereleaseobject (release_id,label,title,version,published_at,content_identity) on/decide,/schema,/explain,/graph,GET /rulebooks[/{id}](current release) and promote-to-live. A pinned release is evaluated from its frozen content — warm or cold caches — after newer releases supersede it; a release of another rulebook/tenant is404with no content; a head naming a release the engine does not hold is a422 rulebook_release_missing— latest is never substituted. Family calls withoutrelease_idresolve the current release when one exists, else the live composition as before (release: null, no backfill). -
Rulebook
/decidereportsruleset_versionasv<release.version>when a release was evaluated (previously alwaysunknown);content_identitykeeps its composition-key format for existing replay consumers. The route now resolves the rulebook once and hands that composition to the navigator (the route/navigator double-read race recorded on #39 is closed). -
Behaviour change on published families: any content edit on a family
that is
activeor already has a release (PATCH title/description/ composition/outcome logic/guidance/scope, POST fields) now resolves and VALIDATES every member live under the owner’s visibility before anything is written; a member that is missing, not live, archived, uncompiled or malformed makes the edit fail with422 release_member_unresolvedand the draft is left unchanged. Previously such an edit saved regardless. The draft is then saved and the release published from it; if that last step fails (422 release_publication_failed) the draft is kept and latest lags it by one release untilactivate(or the next content edit) re-publishes it — latest is always one complete release.PATCH ruleset_refson a promoted (pinned) family now rewrites the pin map (pinned refs only; a floating reference is422 composition_pinned) instead of being silently ignored. Promotion resolves every member of the resulting composition before the candidate goes live. -
Frozen release content is verified against its identity where it is
materialised — the cold compile of a release and the schema/explain/graph
and overlay projections — and a mismatch is
422 rulebook_release_corrupt. Metadata reads (GET /rulebooks[/{id}]) check the release exists and that its identity agrees with the current head, without hashing member bytes;release_errordistinguishes a missing target from a corrupt head tuple. Heads, idempotency tokens and sequence numbers are tenant-scoped, and a release’s sequence number is unique per (tenant, rulebook) by index. -
Rulebook schema responses include
section_names, keyed by section ID. A release freezes each member’s human-authored display name alongside its compiled content, so a pinned application never resolves display text from a newer mutable ruleset document. Releases created before this field exposenullfor names that were not historically captured.
Fixed
GET /rulebooks/{id}/explainreturnedcriteria: []plus anerrorfor every section with rules: the human-readable renderer was handed a model where it walks a dict (#519). Latent since at least 2026-04-26 on an exposed but uncalled endpoint; found by the release read tests.
Fixed
- Rulebook testing publication (#509, #1272). A slug-less candidate can
enter
testingbeside a slug-less current member in the same rulebook slot. Alias-backed replacements retain the existing conflict response.
Fixed
- Preserve every existing member when first promoting a legacy rulebook (#505). Resolve legacy references with the decision loader before any writes, pin floating members to the resolved versions, and reject unresolvable or duplicate references before changing publication state.
- All-green authoring refinement (#507, #1272). Refine mode now asks the authoring model to apply unsatisfied guidance, including metadata-only edits, before its first validation or test call. An already-satisfied request may remain unchanged, while passing goldens alone no longer completes an outstanding guidance edit.
Fixed
- Bounded streamed rule authoring (#125, #583, #1272). Long authoring turns now consume the model provider’s complete streaming message, check for cancellation while it arrives, and preserve existing provider/cache usage accounting. A response that reaches its configured output cap cannot dispatch partial tool input; its retained trace is recorded with a typed, safe output-limit failure for a retry.
- release: verify the Registry’s real response envelope. The official
Registry successfully published
aethis-mcp@0.17.3as active/latest, but the final workflow incorrectly read search results fromentry.nameandentry.versioninstead ofentry.server.nameandentry.server.version, so it reported a false-red after publication. Verification now uses a tested parser with a fixture matching the live Registry response shape.
- release: bind the Registry identity to the case-sensitive GitHub OIDC
namespace. Corrects every current MCP release surface from
io.github.aethis-ai/aethis-mcptoio.github.Aethis-ai/aethis-mcp, the namespace the Registry grants to this repository. A deterministic test now pinspackage.json,server.json, the generated tool inventory, and release verification to that canonical identity.aethis-mcp@0.17.2reached npm but the Registry rejected its lowercase namespace, so it was not a complete dual-registry release.
- release: satisfy the official MCP Registry metadata contract. Shortens
the Registry-facing server description to its 100-character limit and adds a
deterministic test for that constraint.
aethis-mcp@0.17.1was published successfully to npm, but the Registry rejected its overlong description during metadata validation, so it was not a complete dual-registry release.
- release: publish the immutable candidate as a local tarball. The npm
publish step now prefixes the downloaded
release-artefacts/...tarball with./, so npm resolves it as a filesystem package rather than a registry package spec. Thev0.17.0workflow failed before publication; noaethis-mcp@0.17.0package was published.
Fixed
- Cloud Run autoscaling verification. Public candidate gates now accept Cloud Run’s valid revision views where the service maximum is omitted or mirrored exactly, while continuing to prohibit revision-level minimums that would keep retired revisions warm.
Fixed
- Exact Cloud Run build identity. Deployment gates now preserve BuildKit provenance indexes while verifying Cloud Run’s resolved Linux/AMD64 runtime manifest is the unique runnable child of the immutable build image. The same fail-closed identity check protects staging and public production candidates.
Fixed
- Recoverable authoring generation jobs (#481, #434, #420). Generation workers now hold fenced leases with heartbeats and absolute deadlines; abandoned jobs fail explicitly and release their project ownership for an operator-driven retry. Status and cancellation responses expose typed, additive lifecycle telemetry without retaining customer provider keys. Recovery scans are bounded, and operators can pause generation admission while draining old unfenced workers during rollout. Staff job detail uses the same server-authoritative contract version and retry-readiness fields; newer lifecycle records are reported as unsupported and never interpreted as legacy recovery candidates.
- Safe provider failure details for authoring (#481, #463). Authentication, capacity, rate-limit, request and availability failures are classified into stable reason codes while provider response bodies remain server-side.
- Date inputs in generated tests (#481). Strict ISO calendar dates are normalised at the navigator boundary, matching the public decision route, and invalid generated test values fail as contained validation errors.
- Durable public authoring runtime (#481, #485). The public Cloud Run deployment now keeps one instance with CPU available outside requests, limits request concurrency and process-local generation concurrency to one, retains at most one additional local waiter, and declares its CPU, memory, timeout and scaling bounds explicitly. Structured capacity, reaper and admission-pause signals have reviewed alert thresholds. Public staging and production revisions are tagged no-traffic candidates; guarded commands require exact revisions, image digests and database boundaries for drain, legacy recovery, promotion and fence-compatible rollback. Legacy recovery defaults off, and the existing 2 GiB production memory mitigation is preserved independently from staging’s 1 GiB baseline.
Added
-
Compact rulebook choice context for conversational consumers (#477).
POST /decideaccepts opt-ininclude_choice_context: trueand returnschoice_context: complete field-to-section membership plus each authored OR-group’s alternatives, labels, referenced fields and answer-awareopen/closed/satisfiedstatus. This avoids constructing and walking the full dependency graph merely to offer an applicant a choice. The existinggraph_overlayremains available for visualisation and debugging.timing.total_msand decision-log latency now include optional response assembly work, so graph/choice overhead is no longer hidden after the core evaluation boundary. -
Presence-polarity gate at both go-live boundaries (#472; checker from
#448; hardened by the PR #475 adversarial review).
POST /projects/{id}/publishandPOST /rulebooks/{id}/rulesets/{name}/promote-to-liverefuse a ruleset whoseDEFAULT_TO/IS_UNSETis more generous when the field is never asked than when it is answered, with a structured422 non_conservative_presence_opcarrying field, criterion and solver witness. The outcome-level check runs two variants: each presence field freed singly, and all presence fields freed together — the latter catches OR-composed generous defaults that decide a bundle ELIGIBLE with zero questions asked (review CONFIRMED-1). Stated boundary: proper subsets of silent presence fields between those two extremes are not enumerated. An uncompilable stored bundle fails closed as a violation, never a 500. An internal key may passforce_unsafe: true(now also onPromoteRulesetRequest) to proceed — audit-logged (publish_force_bypass_presence_polarity/promote_force_bypass_presence_polarity). Per the owner decision of 2026-08-23, a republish of an existing reserved-namespace (aethis/*) target — a prior holder of the slug (including the bundle’s own stored slug on a slug-less republish) or a live rulebook member with the sameruleset_name— is ADVISORY instead of blocking: the response carriespresence_polarity_advisoriesand the advisory is audit-logged too (publish_presence_polarity_advisory/promote_presence_polarity_advisory), becauseaethis/construction-all-risks’s two accepted defaults are load-bearing for published benchmark artefacts (Aethis-ai/aethis-examples-internal#38). A brand-new first-party slug or member name binds like anyone else. The advisory/binding decision is one named function (presence_polarity_gate_mode), so retiring the carve-out is a one-line, cited change.
Added
- Typed recovery support for asynchronous ruleset generation (aethis-core#481):
Aethis.get_generation_status()/AsyncAethis.get_generation_status()now returnGenerationStatusResponse, including structured job progress, worker-heartbeat, lease/deadline, test, and terminal-failure telemetry. Aethis.cancel_generation()/AsyncAethis.cancel_generation()explicitly request cancellation of an observed generation job by project and job id, preventing a delayed request from cancelling a successor job, and returnGenerationCancellationResponse. Cancellation is cooperative: it fences the job and releases the project, while an in-flight provider request stops at its next safe boundary. The SDK never cancels, resumes, retries, or stores credentials on a caller’s behalf. Requires the aethis-core generation- recovery API to be live on the target API; earlier engine versions do not advertisegeneration_contract_version=1and are refused before mutation. Status exposes telemetry availability, server-authoritative worker lifecycle, and retry readiness; cancellation distinguishescancelledfrom the idempotentalready_cancelledoutcome.
- feat: generation status and cancellation. Adds
aethis_generation_statusto inspect a project’s current or most recent authoring job, andaethis_cancel_generationto abandon a target-bound observed job (project id plus matching job/confirmation ids) and release project ownership. Worker shutdown may be cooperative rather than immediate. Status is read-only; cancellation is explicitly annotated as a destructive API-key mutation, so MCP hosts can gate it for approval. Matching ids prevent retargeting but do not prove human consent; agents must obtain a fresh explicit reply before calling. Both responses and diagnostics remain fenced as untrusted API data. Status carries telemetry availability, server-authoritative worker lifecycle, and retry readiness; cancellation preserves the idempotentcancelled/already_cancelledoutcome. - safer authoring timeout guidance. A timed-out
aethis_generate_and_testnow tells agents to inspect status before retrying and to cancel only when the caller asks to stop the run. This release requires aethis-core’s project status and generation-cancel endpoints to be live.
- feat: explicitly recover an authoring project from an abandoned generation.
aethis cancel [-p PROJECT]first observes and displays the exact active job, then calls the engine’s job-bound cancellation endpoint after a confirmation (--yesfor automation), marks that job failed, and releases its ownership of the project. It reports the engine’s honest limitation: an already-running worker may continue even though a new run can now be admitted. The idempotent response distinguishescancelledfromalready_cancelled. Polling never cancels automatically. - feat: generation status shows live convergence and heartbeat telemetry.
aethis statusand the livegenerateprogress line consume the engine’s turn count, best test pass rate, tool count, last tool, andseconds_since_progress. Older engines remain compatible: absent telemetry is omitted rather than guessed. A client-side timeout now names the exact status and cancel recovery commands. Status also renders telemetry availability, server-authoritative worker lifecycle, and retry readiness.
- fix(generate): a dropped connection mid-poll no longer strands the run on a
stale ruleset id.
--pollguards against an interrupted generation leaving.aethis/state.jsonnaming an earlier ruleset, but the guard only coveredAethisAPIError. A dropped connection, DNS failure, TLS error or read timeout raiseshttpx.HTTPError, which unwound past it — so the failure most likely to interrupt a long poll was the one that bypassed the guard, and every laterdecide,test,explainandfields pullsilently answered from the wrong ruleset. The poll now absorbs a short connection blip and finishes normally; where the API is genuinely unreachable it names the stale id and prints the command that recovers the real one. - fix(generate): the stale-pointer guard can no longer fire on a run that succeeded. The pinned-vs-produced field diff runs after the new id is recorded, so a connection lost during it made the CLI announce “Done! Ruleset: X” and then insist X was from an earlier generation — a false claim about the exact fact the guard exists to keep honest. The diff now honours its own “never fails the command” contract for transport errors, and the guard never describes a pointer this run recorded as stale.
- fix(generate): a blip while waiting for the new ruleset id no longer discards it. After the engine reports success the CLI re-polls for the id. A transport failure there stranded a generation the CLI already knew had succeeded; those attempts are now individually tolerant.
- fix(generate): a persistent outage is reported as unreachable, not as a timeout. The retry is bounded by consecutive failures, so a dead connection surfaces as one rather than being absorbed to the deadline and reported as a slow job. A flapping link is deliberately not aborted — it is working, and killing it one poll short of success would deliver the very failure above — so a poll that suffered dropped connections says so when it does time out.
- fix(generate): a failed publish no longer reads as a failed generation, and
says why. A transport error on the auto-publish reported only “Could not
reach the Aethis API”, so a successful generation looked like a failure and
got re-run. It now reports the ruleset, names the reason the publish did not
run, and names
aethis publishas the step to re-run — previously the API-error half of this path gave a green tick with no reason at all. - fix(generate): a success that never surfaces an id says so. That ending also leaves the recorded pointer naming an earlier generation, and was silent.
- fix(generate): an unreadable value space is reported, not raised. The
field diff verifies value-space-pinned fields against the registry and
formatted the failure as if it were always an API error, so a dropped
connection there raised
AttributeError— an unhandled traceback on an otherwise successful generation. - fix: an unreachable API cannot crash the error reporter.
httpxexposes.requestas a property that raises when unbound, so the defensivegetattr(e, "request", None)in the transport-error path could raiseRuntimeError— a traceback about the error handler in place of the one actionable line. The rendering now has a single home shared by the top-level boundary and the commands that must clean up before exiting.generateandrefinename the host they were actually configured with rather than the public default; other commands still report via the top-level boundary, which readsAETHIS_BASE_URL.
Measured against the deployed
0.57.0: additive on the public wire (owner
probes of an exact private draft via /decide, two opaque per-field metadata
keys on both schema routes, durable authoring-source identity) plus one
navigation correction that changes which question /decide asks — never
the verdict — on rulesets with alternative Boolean branches.Fixed
-
A Boolean route is no longer flattened into its field union — dead
branches stop driving questions (aethis-core#465, #466; defect shape
DS-63; workspace epic ws#1067). For a criterion such as
(kind = ordinary AND taught) OR (kind = certificate AND aquals AND elps)the compiler always produced the right condition, but the navigator also kept the flat union{kind, taught, aquals, elps}, and twelve consumers —find_next_input_field,relevant_unanswered_fields,_check_requirement_status,get_best_requirement,get_optimal_path_to_eligibility,_add_input_field_constraintsand its tie-breaks, the layered explanation’sunused_facts,routing.py, the scenario evaluator,build_pydantic_model,detect_review_triggers— treated that union as though every member were conjunctively required. Withkind = certificatethe engine could still asktaught, report it missing, cost it on the optimal path, or hold the requirement atunknown. EachEligibilityRequirementnow carries one requirement-scoped, linear NNF semantic plan built from its unresolved AST (NOTpushed to literals,IMPLIES(a,b)lowered toOR(NOT a, b), no DNF materialisation, no route-count fallback). Decision/status evaluate it under_resolve_presence_ops; askability, selection and relevance evaluate the same plan under_relax_presence_ops, so an unansweredDEFAULT_TO/IS_UNSETfact that could still change the outcome stays askable. Duplicate criterion ids across composed sections keep requirement-scoped identity and collision-free solver selectors. Non-applicant fields remain selectable at a high finite cost so an applicant-completable branch wins when one exists. The flatinput_fieldsunion survives as provenance/schema metadata only; if the engine cannot attribute the authored structure it fails loudly rather than asking the union. Verdicts are unchanged on every existing rulebook (the DSL parser and truth evaluation were never shown to be wrong; read-only scan of all 257Rules_Stagingbundles: 0 flat bundles with authoredoutcome_logic). Visible correction:status_report["path"]may now name the actually satisfied requirement where the old flat completion proxy returnedNone. Reproducer: the private English-language draft askedeng.degree_taught_in_englishon the Australian postgraduate-certificate route (6/7 question-order probes onmain; 7/7 after). A narrowmissing_fields-only over-report on two flat-bundle shapes is tracked in #467; decision, next question and optimal path are correct there too. - Explicit field pins are authoritative through the single post-parse normalisation seam (#469, #470): authored label, question, weight, owner, injection, recovery and UI metadata on a pin survive fresh generation and refinement alike, absent-vs-explicit-null is preserved at ingestion, and the same path serves inline and named value-space fields.
-
Field pins are enforced inside the authoring TDD loop, not only at
persistence (#460, #461):
compile_and_testappliesnormalise_parsed_fieldsbefore compiling, so a missing/unparseable pin or invalid space expansion reaches the model during iteration instead of killing the job at the end. - Authoring normalisation failures keep their reason code, affected fields and parse warnings in the job error detail (#458, #459), and the authoring preflight’s exact token count is emitted through the deployed application logger (#456, #457) — observable on staging without logging prompt or source text.
Added
-
Presence-operator polarity checker — report-only (#448; refs
tda-server#1351).
find_non_conservative_presence_ops(per criterion) andfind_outcome_polarity_violations(per composed outcome) detect aDEFAULT_TO/IS_UNSETwhose never-asked reading is more favourable to the applicant than answering it would be — a satisfying model ofC[all defaulted] AND NOT C[f free]is a concrete world in which declining to askfpasses while some real answer fails. Both readings come from the navigator’s own_resolve_presence_ops/_refresh_conditions_from_answers, so the check cannot drift from the semantics it certifies; it fails closed. Measured:form-an-english-language-r2has no presence ops;form-an-life-uk-r2andaethis/consumer-credit-prequalificationare clean;aethis/construction-all-riskshas 2 violations (verdict flips with ask order, witness-driven through the shipped navigator). Nothing calls it yet — binding it would break that published ruleset on contact (adaptive-gating.mdrule 5); the measurement above is what the gate decision now has. -
Authenticated ruleset owners can evaluate an exact private draft through
/decidewithout publishing or activating it (aethis-core#464; workspace epic ws#1067). The fallback accepts only the immutable compiled-ruleset id and the owning tenant: draft slugs, project aliases, anonymous callers and other tenants still receive 404, and the request remains read-only. This gives authoring acceptance harnesses a supported way to test question order on the exact generated artefact before it can become live. -
Rule-authoring inputs now have durable identity and exact preflight
diagnostics (aethis-core#452; workspace epic ws#1081). Identical concurrent
uploads within one tenant/project/filename reuse one stored source identity
and report
newversusreusedcounts; inactive matches and changed active filenames return actionable PATCH-lifecycle conflicts. New generated citations use immutablesource_id#item_idkeys, and source hard deletion is disabled in favour of soft supersession/reference-only status. Both generation routes assemble one immutable prompt, make one authoritative provider token-count call, and execute that same prompt object. Rejections report the exact total plus deterministic UTF-8 byte/character sizes by component without claiming per-component token attribution. Empty eligible corpora and unresolved generated citations fail before ruleset persistence; concurrent requests are admitted through the existing project state.
Changed
- Internal production engine scales to zero; staging stays warm (#462).
The
$_ENVIRONMENTbranch incloudbuild.yamlnow setsmin-instancesper environment; part of the August cost response (ws#503).
- fix(generate): authored field behaviour is deterministic. Generation uploads now carry explicitly authored labels, questions, ownership, injection source/phase, ordering weight, recovery capabilities, and UI hints on the field pin. The engine’s normalisation seam can therefore preserve the authoring contract instead of accepting model-invented replacements.
- fix(generate): an omitted question on a non-applicant field is explicit.
When a field declares a caseworker, system, or derived owner and has no
question, the CLI sendsquestion: null; this clears any question invented by the generation model. Legacy fields that declare none of the new metadata retain their previous upload shape. - fix(generate): capability checks cover behavioural metadata. An older engine that would silently discard any declared property is refused before the field-spec push, just as it already was for enum labels and canonical storage mappings.
- fix(generate): source upload results are truthful. Generation reports the engine’s exact new and reused counts instead of calling every attempted file newly uploaded.
- fix(generate): edited sources remain regenerable. Before uploading changed content under an existing filename, the CLI supersedes the prior active source, uploads and links its replacement, and reactivates the prior source if the replacement upload fails. This avoids both a hard-blocked edit loop and duplicate active source text in the authoring prompt.
Added
-
SchemaFieldreads the two authored metadata properties the engine now publishes (aethis-core#449; workspace epic aethis-workspace#1067).enum_labels— an optional{member slug: display label}map — andcanonical_field— the author’s pairing between an eligibility field key and the consumer’s canonical storage key — are carried on every field of both/rulesets/{id}/schemaand/rulebooks/{id}/schema. Previously the SDK’s model declared neither, so Pydantic silently dropped them and an SDK caller could not render a label for a stored enum slug. Both are opaque to the engine and to this SDK: nothing here validatesenum_labelsagainstenum_values, or resolvescanonical_field. Both are optional and default toNone, so a schema from an engine predating the properties parses exactly as before.Noneand{}are deliberately distinct —{}is an author declaring “no labels”,Noneis an author declaring nothing — so read them withis None, never truthiness.
A field can now say what its members are called, and which stored key its
value belongs to. Requires an engine that models both on a field spec —
one that does not is refused before the upload, never allowed to accept it
and drop them. Projects declaring neither are unaffected.
- feat(fields):
enum_labels:on a field pin. An enum field may carry a per-member map of display wording —enum_labels: {ion_drive: Ion drive}— beside the members themselves. The engine treats the labels as opaque text and publishes them on both the ruleset and the composed-rulebook schema, so every consumer renders the same wording from one authored source instead of each keeping its own copy of it. Optional per member: label the ones whose slug does not read well and leave the rest. Both properties are handled on presence, not emptiness: an explicitenum_labels: {}is a declaration the engine keeps and republishes as{}, so it is transmitted, validated and preserved across afields pullrather than being discarded as an empty value — and it fires the capability check below like any other declaration. - feat(fields):
canonical_field:on a field pin. The storage key a field’s value belongs to (canonical_field: spacecraft.propulsion) is authored beside the field and published on the schema, rather than re-derived from the key by every consumer that needs the pairing. Any field type may declare it. - feat(fields): both are validated locally before they can reach the engine.
enum_labelsmust be a mapping of member → non-empty text on an enum field, and where the members are declared inline it may not label a member the field does not have — the engine accepts such a label and simply never renders it, which is invisible at every layer afterwards. Where the members are a namedvalue_spacethey live on the registry, so only the shape is checked.canonical_fieldmust be non-empty text. - feat(generate): an engine that cannot keep them is named, not written to. A field-spec property an engine does not model is ignored rather than rejected, so the upload succeeds, the generation runs, and the authored values are gone — the same silent-drop shape the value-space pin was guarded against. The CLI reads the engine’s own published schema and refuses before the push, naming the properties and the engine. An engine whose schema could not be read at all is a different answer and says so: the upload proceeds with a warning to check the published schema afterwards, because unknown is not unsupported.
- Unchanged for everything else. A project declaring neither key produces the byte-identical upload payload it did before, and never triggers the capability probe — so it keeps working against any engine, of any vintage.
- feat(rulebooks):
set-fieldsis guarded the same way.aethis rulebooks set-fieldsposts its fields file as authored, so it already carried both properties — and would equally have had them silently discarded by an engine that does not model them. It now refuses first, asking about the rulebook field model rather than the project one, since an engine can carry the properties on one and not the other. A fields file declaring neither key is untouched and never probes. That command still performs no other validation of its file (#114).
Tagged 2026-08-19 without a changelog entry; documented here after the fact.
Additive on the public wire.
Added
-
next_questionsays which section(s) the question serves (#446).NextQuestion.sections: List[str](default[], so older callers see no change) lists every composed section whose requirements read the selected field — a set, not a single value, because the first question of a multi-section rulebook is typically a shared fact. Populated onnext_questionand on theoptimal_pathfallback; derived in place fromsection_to_groupswith no extra solve, and fails soft to[]so a narration aid can never cost a decision. Also exposesEligibilityNavigator.sections_for_field(name). -
Two pieces of authored per-field metadata now ride the published schema
(aethis-core#449; workspace epic ws#1067).
enum_labels— an optional{member slug: display label}map — andcanonical_field— the authored pairing between an eligibility field key and the consumer’s canonical storage key — are carried from authoring to both/rulesets/{id}/schemaand/rulebooks/{id}/schema, throughExpectedFieldSpecingestion, the DSL emit/parse round trip, the compiledFieldDefinition, and the rulebook field vocabulary (where a declaration overrides the member ruleset’s, likequestion). Both are opaque to the engine: nothing compiles, solves, gates or validates them, andenum_labelsis deliberately not checked againstenum_values— a referenced field’s members come from the value-space registry and may move independently of its labels. They exist so a consumer can render a human label from the single stored slug instead of keeping a second copy of the vocabulary; a hand-maintained copy of the Form AN enum drifted from the authored source and silently discarded an applicant’s answer, which is what this deletes. The pin is the sole origin. A pin that declares a property sets it; a pin that is silent CLEARS it; a field with no pin cannot carry either. So a property can be removed by authoring, not only added, and a model cannot invent one. For a field shared by several rulesets in a rulebook, the first member in the rulebook’s pinned order that DECLARES a property supplies it — which is not the same as the first member that merely mentions the field. Additive and optional throughout. A build that omits them persists neither key, on both of the populations that store them — the compiled field (they join_CONDITIONAL_FIELD_KEYS) and the rulebook’s own field vocabulary (guarded at both the Pydantic and the BSON layer, because Beanie encodes nested models without callingmodel_dump(), and every rulebook write is a full-document save). So a field authored without them is byte-identical to today and a rolled-back engine still re-validates it. Both schema routes emit an explicitnullso a consumer can tell “no labels authored” from “this engine predates labels”.
Measured against the deployed
0.49.2 on 2026-08-17, this release is purely
additive on the public wire: three new routes, nine models gaining fields, four
new schemas, and nothing removed or renamed. No consumer change is required
to keep working.Added
- Named value spaces (aethis-core#423/#428, #424/#429, #426/#432; workspace
epic ws#980). A model may now reference a shared, named vocabulary instead of
retyping the same enum on every field.
GET/POST /api/v1/public/value-spacesandGET/PUT /api/v1/public/value-spaces/{name}expose the registry;ExpectedFieldSpecgainsvalue_space; expansion happens deterministically at compile time, so a decision is never at the mercy of registry state at solve time. Ships with drift telemetry over the generation record. - Engine-level guidance tier (#427/#430).
GuidanceListItemandAddGuidanceRequestgaintier,overrides,adherence,principle_keyandprocess_type, with admin routes at/api/v1/admin/guidance/engine. Adherence is reported honestly rather than assumed. - A rulebook declares whether its outcome is a COMPLETE determination (#412).
complete_determinationonCreateRulebookRequest,UpdateRulebookRequest,RulebookResponseandDecideResponse, so a caller can tell “this rulebook has decided everything it governs” from “this is all it can say”. - Typed
undetermined_reasonon/decide(#402, ws#884). The positive terminality signal: consumers no longer have to infer why a case is undetermined from the shape of what came back. pending_non_applicant_input, pluselicitation_ownerandrecoverable_fromonNextQuestion— who may supply a given fact is now part of the wire contract rather than something a caller reconstructs.POST /api/v1/public/projects/{project_id}/generate/canceland live turn/convergence telemetry during authoring (#403).- Replace semantics for
POST /tests(#404) — adding a suite no longer silently duplicates it. - A production deploy gate (#441, closes #442). A tag previously ran no
tests at all:
test.ymlhas notags:trigger andcloudbuild.yamlhas no test step.scripts/preflight_production.pynow blocks a tag unless the tree is clean onmain, CI is green on that exact commit, staging is verifiably running the code being promoted, and a behavioural smoke passes against it;.github/workflows/post-deploy-verify.ymlre-runs that smoke against prod after the roll and opens an issue if the deployed engine stops deciding correctly. The oracle lives inscripts/smoke_expectations.json.
Fixed
- A group dead only by a presence-default keeps offering the reviving field (#418) — the navigator no longer closes a path the applicant could still open.
- The final best-source re-evaluation carries the value-space context (#435), so a field resolved against a named space is not re-scored without it.
- An authenticated owner may read their own draft ruleset’s
schema/explain/graph(#436). Ownership, not publication status, governs the read. - Deploy-time database guard (#408/#409):
DATABASEis derived from_ENVIRONMENT, and a prod deploy against a staging database is refused outright rather than discovered afterwards. RULES_AUTHORING_MODELhas a durable, staging-only home (#405/#406) —--set-env-varsreplaces the whole env, so a hand-set value was wiped by the next deploy.- Ruff rule set pinned via explicit
select(#415); the 0.16 default broadening had turned the daily dependency heartbeat red.
Changed
- Model catalogue refreshed to the gpt-5.6 family (#401), repointing the mini class like for like.
A member section’s field pin is its own contract, not the rulebook’s.
- fix: rulebook-level fields no longer extend a section’s pinned field spec. The merge pinned every rulebook field onto every member section, demanding fields the section’s rules never author — and the engine historically dropped them silently (every published
referees_identitybuild was missing 3-5 of the rulebook’s cross-section fields). With the engine’s new loud pin-presence gate, such a generation now correctly fails — which surfaced the over-broad pin on the first live canary run of the named-value-spaces epic. The rulebook’s definition still wins on keys the section declares itself (the canonical definition is unchanged); it just never adds keys. The drift report’s comparison universe changes identically, so it now reports on exactly the section’s own contract.
Named value spaces reach the wire (aethis-core#424, epic aethis-workspace#980 P3). Requires an engine exposing
/api/v1/public/value-spaces for referenced fields; projects with no value_space: pins are unaffected.- feat(fields):
value_space:on a field pin.fields.yamlenum fields may reference a named, versioned value space (e.g.value_space: form-an/countries) instead of inliningenum_values— the engine expands the members deterministically and the model never authors them. The two are mutually exclusive per field, validated locally; reference-form enums no longer fail the “enum needs enum_values” check. - feat(generate): registry sync before spec-set. Locally-authored space files (
value_spaces/orshared/value_spaces/*.yaml, wire-formname/members/provenance) are PUT to the engine registry before the field spec is pushed, carrying the sync-state base version so a stale checkout gets the engine’s 409 (“space moved under you — pull first”) instead of silently regressing the shared vocabulary. Any non-2xx aborts before spec-set naming the missing engine capability — an older engine ignores unknown pin keys and would otherwise silently drop the reference. - feat(generate): registry-aware drift report. A referenced field is verified against the registry at the exact version the generation resolved (from the job result), printing
<key> ← <space>@vN (M members verified)— never the old “pins no enum_values ⇒ opts out” silence on exactly the migrated fields. A space that advanced between sync and generation is flagged. An unverifiable referenced field is reported, never skipped. - feat(fields):
pullpreserves the migration. A local field declaringvalue_space:keeps the reference; the server’s expanded members are never written back intofields.yaml.
The releases start reaching PyPI again.
- fix(packaging): explicit package discovery. Every release from v0.30.0 to v0.34.0 was tagged and never published: the publish workflow’s own integrity step creates a top-level
evidence/directory before building, and setuptools flat-layout auto-discovery aborts the build the moment any second top-level directory exists (#94). Discovery is now explicit (include = ["aethis_cli*"]), so a stray directory — the workflow’s or anyone else’s — can no longer poison the build. Proven both ways: withevidence/present the build failed before this change and succeeds after it. This release supersedes v0.30.0–v0.34.0, none of which reached PyPI; it is the first PyPI release to carry everything since v0.29.0, including--no-publishand the member-set drift report.
Generating a ruleset no longer has to activate it.
- feat:
aethis generate --no-publishleaves the ruleset an unpublished draft. A successfulaethis generate --pollpublishes what it produced, and publishing activates it. For authoring that must leave a draft behind — a ruleset that only ever activates somewhere else, after promotion — there was no way to opt out, and the only workaround was to archive it immediately afterwards, which leaves a real window where it is live. The flag suppresses the publish and nothing else: the poll, its timeout, and the post-generation field diff are untouched, because those are what make the run worth doing. - feat: the ending says which one it was. A run that was told not to publish reads differently from one whose publish failed. Both leave a draft and both point at
aethis publish, so collapsing them would tell an author who never passed the flag that everything went to plan — precisely when they most need to know it did not. The deliberate ending names the flag; the failure ending is unchanged. - Unchanged without the flag. It defaults off, and
aethis refine, which shares the same machinery, was not asked to change and does not: the publish still happens on the same call with the same argument.
A pinned enum’s members are checked, not just its name — and the check survives
the failure it was written to explain.
- fix: a failed generation now prints the drift report before it exits. The diff existed for the case where the model did not honour the pin, and that was the one case it could not reach: the poll loop exited the moment a job reported
failed, several frames below the call that would have printed it. A failed generation said “Generation failed” and nothing about what had actually been produced. The poll now reports how the run ended and the caller decides, so the diagnostic runs first and the exit code is unchanged. - fix: the diff is computed against the ruleset this run produced — never the last one recorded. It previously read the ruleset id off
.aethis/state.json, which is written only on success, so a run that produced nothing would have been diffed against an earlier generation and the result would have read exactly like this one’s. Where a run names no artefact — which is every failure today, since the engine records one only on its success paths — the CLI says so plainly instead of falling back on anything. - fix: the ruleset a successful generation reports is the one that job produced, not the project’s newest. The id was read from
latest_ruleset_id, which is a property of the project — so with two generations running against one project it could name another run’s artefact, both in the diff and in the id written to disk for every later command to default to. The job’s ownresult_ruleset_idis now preferred, with the project-level value kept only as a fallback for engines that record nothing on the job. - fix: a generation whose polling breaks off no longer goes quiet about the recorded ruleset. An API error mid-poll — a 500, a dropped connection — exits without a verdict, and the job may well have succeeded regardless. Like a timeout, nothing has been ruled, so the recorded id stands; unlike before, it is named, because that ending was otherwise the remaining way to reach a silent
aethis fields pullfrom an earlier generation. - fix: a failed generation no longer leaves a ruleset pointer that reads like its result. The recorded id is what
aethis fields pulldefaults to, so after a failure the next pull synced from the previous generation without a word, and the fields appeared to be the ones just generated. A failure now clears it — naming it, so it can still be passed with--ruleset-id— andfields pullrefuses rather than guessing. A timeout is treated differently on purpose: nothing has been ruled, the job may still land, so the pointer stands and is named as stale instead of discarded. - fix: an enum that came back with no members is reported as drift, not success. The member comparison skipped a produced enum whose member set was empty, reading it as “nothing to compare” — so a field whose members had all been dropped printed
✓ Fields: all N pinned field(s) were produced. That is the worst case the report exists to catch, reported as the best one. The empty set is now the loudest result, and a field that came back as something other than an enum reads the same way. - fix: a schema that cannot be read is said out loud. The fetch failure was swallowed entirely, so a purged or unreachable draft produced a run with no verdict at all — indistinguishable from a clean one. It is now reported, naming the ruleset, and still never fails the command.
- fix:
aethis generatecompares pinned enum member sets, not only field keys. The post-generation drift report asked whether each pinned field was produced. It was — so a schema whose enum had quietly grown a member nobody pinned printed✓ Fields: all N pinned field(s) were produced. The members were already in hand locally; only the comparison was lossy. It now checks each pinned enum by exact equality in both directions and names the added or dropped members. - Why it matters beyond tidiness. A field can be present, correctly typed, and still wrong. Where an enum is used as an escape hatch — a value that keeps a decision at undetermined pending human review — the safety property is that no value the schema allows can turn that into a definitive “no”. One unpinned member breaks it, and no test case can catch that: a test occupying the offending value would itself be the bug. Five consecutive generations of a real ruleset were scored as clean this way.
Re-uploading a test suite no longer leaves a second copy of it.
- fix:
aethis generatereplaces the project’s test cases instead of adding another copy of them.tests/scenarios.yamlis the authoritative suite and it was uploaded in full on every run, while the API only ever appended — so N authoring cycles left N copies of every case. Nothing errored: duplicates inflate the denominator of every pass rate, so a run reported a total that looked like a result and was partly copies of itself. Two real projects were found holding 185 and 96 cases against files of 42 and 14. The upload now asks the engine to replace, which makes it idempotent. - feat: the count of cases the upload overwrote is reported. Replacing is destructive — it removes what was on the project — so
aethis generateprints how many cases went and how many arrived (Uploaded 42 test case(s) from scenarios.yaml — 42 replaced) rather than performing it silently. - feat: an engine that cannot replace is named, not worked around quietly. Replacing needs an engine that offers it, and one that does not offer it does not reject the request — it ignores the member and appends, so sending it blindly would restore the duplication with nothing to notice. The CLI reads the engine’s own published schema, sends the flag only where it is advertised, and where it is not — or where the schema could not be read at all, which is a different answer and says so — prints that the upload APPENDED, that running again adds another copy, and what to do about it. Nothing is dropped silently in either direction.
Fixed
-
A rulebook’s field vocabulary now reaches the navigator
/decidesolves against, not only the/schemaprojection (aethis-core#396). The rulebook is the surface on which a fact’s ownership is corrected without regenerating a member ruleset.0.53.0implemented that correction in exactly one place — theGET /rulebooks/{id}/schemaroute — and nowhere on the solve, so the two public surfaces disagreed about the same field in the same request window. Measured on the live engine:/schemareportedelicitation_owner: caseworker, while/decideproposed that same field to the applicant asnext_question, labelled itapplicant, and returnedpending_non_applicant_input: null. Ownership itself was never broken — the engine honours it exactly whenever the compiled field carries it. The overlay simply never reached the navigator, which is compiled from the members.-
elicitation_overlay()(public/models/rulebook.py) is now the single definition of what a rulebook’s vocabulary says about a field’s four elicitation axes./schema, the navigator and/decide’s field registry all read it; none re-implements the merge./schema’s output is unchanged. -
The overlay is a LEFT JOIN applied to the compiled navigator, and it can
restrict askability but never grant it — so a vocabulary entry
added purely to reword a question cannot silently re-open a
caseworkerfact to the applicant. The effective owner is the union of “the compiled field says non-applicant” and “the rulebook says non-applicant”; anelicitation_ownerofapplicant, and an emptyrecoverable_from, contribute nothing. That union is not a stylistic choice — it is the only merge that is safe on data written by0.53.0, which materialisedelicitation_owner: "applicant"andrecoverable_from: []into every stored spec. No later read can distinguish those bytes from a deliberate declaration (workspace defect-shape DS-43), so any “the vocabulary wins where declared” rule downgrades a compiledcaseworkerfield on every legacy row. Under a union those bytes are inert whatever they meant.injection_sourceandinjection_phasewere Optional in0.53.0too, so for those absence is observable and they useis not None. The cost is that a rulebook cannot re-open a compiled non-applicant field to the applicant. Measured before adopting it: 0 of 352 compiled fields declare a non-applicant owner, so the capability has no current use case, and the asymmetry runs the safe way — a withheld question surfaces as a visiblepending_non_applicant_inputentry, whereas a wrongly-asked one is the defect this whole contract exists to prevent. It changes who may supply a fact, never whether the fact constrains the solve — the omitted field still decides the case when its value arrives. -
Corrections work in both directions: a rulebook may also declare a
compiled
caseworkerfield applicant-owned, and the engine will then ask it.
-
-
pending_non_applicant_inputno longer degrades silently. The projection swallowed every exception into[], which is indistinguishable on the wire from “nothing outstanding” — so a broken projection reported every waiting case as finished, the exact failure the field exists to prevent. The documented tolerance (a navigator predating the projection) is kept but narrowed toAttributeError; anything else is logged at ERROR with a traceback. Either way the caller is told — and a degraded body is never written to the decision cache, because the header that marks it degraded is a per-response artefact that does not ride the cached payload. Caching one would have served “nothing outstanding” with nothing to contradict it for the full TTL (24h for a pinned rulebook) — the same silent degrade, one layer down. Healthy bodies are cached as before.
Added
X-Aethis-Pending-Inputresponse header on/decide. Present only when the outstanding-non-applicant-facts projection could not be computed (unavailable; reason=unsupported|error|malformed). Absent on every healthy response. Consumers that treatpending_non_applicant_input: nullas “nothing outstanding” should treat a response carrying this header as “unknown” instead. Mirrors the existingX-Aethis-Records-Omittedheader on the catalogue route.
Changed
-
RulebookFieldSpec’s four elicitation axes are nowOptional, defaulting toNoneinstead of to"applicant"/[].Nonemeans the author did not declare one, and the compiled field’s value stands; it is not a new owner class.GET /rulebooks/{id}/fieldstherefore returnsnullfor an undeclared axis on documents written from0.54.0onward. Rows already in the database were written by the0.53.0model and still carry their materialised"applicant"/[]; there is no migration, and none is needed, because the overlay’s union rule (above) makes those values inert.GET /schemastill emits an explicit"applicant"when neither the vocabulary nor the compiled field declares an owner. -
GET /rulebooks/{id}/schemachanges its answer for some existing stored documents. The0.53.0route treated any vocabulary entry as authoritative for all four axes;0.54.0treats a storedinjection_source: null/injection_phase: nullas absence and falls through to the compiled field, and ignores a vocabularyelicitation_owner: "applicant"/recoverable_from: []entirely. A differential over stored-document shapes finds 350 of 512 stored-row combinations diverge — e.g. a vocabulary entry storing nulls over a compiled definition that declares a source and phase now reports the compiled values rather thanNone. Every divergence is in one direction: the compiled declaration is no longer silently erased by a materialised default, and no field becomes more askable than0.53.0reported it. That direction is asserted as a property test, not just observed. On the currently deployed vocabulary the served bytes are unchanged, which is why this was initially and wrongly recorded as “byte-identical”: 0 of 352 compiled fields declareinjection_source/injection_phase/recoverable_from, and all 3 stored vocabulary entries declareelicitation_owner: caseworker, so none of the 350 shapes occurs in production today. That is a fact about the data, not about the code. -
content_identityfor a rulebook decision binds the rulebook’s field vocabulary — but only when the rulebook actually carries corrections. The vocabulary determines what the engine asks and what it reports as pending, so a decision taken under one vocabulary is not replayable against another, and an identity ignoring it claimed a replayability it did not have. A rulebook with no corrections keeps its0.53.0identity byte-for-byte (the component is appended only when present, never as a:nonesentinel): nothing about how it decides has changed, and consumers compare recorded against replayed identities to decide whether to serve a decision explanation — churning every identity uniformly would report “the rules have been updated” for rulebooks where nothing had. A rulebook that does carry corrections gets a new identity, which is the correct answer. The decision cache and the navigator cache use the same key, so aPOST /rulebooks/{id}/fieldsedit now correctly invalidates both — the vocabulary is the one key component that is mutable in place, with no version cut, so before this a correction would have been masked by a warm entry for the whole TTL.
Added
-
injection_phasegainspost_submission(aethis-core#397). The lifecycle vocabulary ended atpre_submission, but some facts are produced by an assessor reviewing the evidence submitted with the application — they cannot exist before the submission that triggers them. Authoring such a field had no honest option: omitting the phase loses the due-date the axis exists to provide, andpre_submissionasserts the fact is due before the thing that produces it.POST /rulebooks/{id}/fieldsrejected the authored value with a 422. Purely additive: the vocabulary stays closed (an unknown phase is still aValidationErrorat authoring time andunknown_injection_phaseat publish), every previously-valid value remains valid, and no stored content changes meaning. Consumers treat the phase as an opaque reported label — nothing compares, sorts or indexes it — so the new value needs no ordering decision.
Changed
-
INJECTION_PHASESandELICITATION_OWNERSare removed. Each vocabulary existed twice — as aLiteral(which rejects a bad value at authoring time) and as a hand-maintained tuple (which the publish gate renders into its error) — so adding a value to one and not the other produced a value that models accept and the gate rejects. The publish gate now derives each vocabulary from its type at the point of use, so the second name does not exist to fall out of step. No behaviour change. Deriving the tuples rather than deleting them was considered and rejected: an equality assertion catches the two copies diverging, but a contributor who reverts the derivation and re-types the values correctly ships green, leaving a hand-maintained copy for the next value addition to split. Internal constants only — both names were consumed solely byaethis_core.public.validation.field_elicitation, and neither appears in any request or response shape. -
ElicitationViolationcarries the vocabulary it applied, as a newallowed: Tuple[str, ...]field (empty when no closed vocabulary applies). Violation messages are byte-identical to before — verified across every violation case — and the routes that consume violations readfield_keyandmessageas they always did; this adds a structured reading of what the prose already said, so the guard for it need not parse prose back out.
Added
-
Fact ownership and document recovery are now two independent, typed axes on
every compiled field, and
/decidereports what it withheld (aethis-core#389, workspace#835).0.52.0putelicitation_owneron the rulebook vocabulary; this puts both axes onFieldDefinitionitself, so a ruleset carries them without a rulebook overlay, and makes the engine act on them rather than merely serve them.elicitation_owner(applicant|caseworker|system|derived, defaultapplicant) — WHO may supply the fact.recoverable_from(list of evidence-capability names, default empty) — WHAT could supply it without asking. Independent of ownership: a date of birth is applicant-elicitable and passport-recoverable at the same time, so collapsing the two would either silence a good question or block a good document. Absence of a friendly question is explicitly not the safety mechanism.injection_source/injection_phase— for a non-applicant fact, the named mechanism expected to supply it and the lifecycle point at which it is due.
GET /rulesets/{id}/schema,GET /rulebooks/{id}/schema,next_questionand everyoptimal_pathentry, and are always emitted (including theapplicantcase) so a consumer reads askability instead of inferring it from an absent key. -
pending_non_applicant_inputon/decide— typed{field_id, owner, source, phase}for every required fact the applicant is not the source of which has not been injected. This is the half that stops the omission becoming a lie: with no applicant question left, anundetermineddecision and a nullnext_questionis otherwise indistinguishable on the wire from a finished case. The engine now selects only applicant-owned fields, and omits the rest fromnext_question,optimal_pathandmissing_fields— but they still constrain the solve, so the outcome stays undetermined until the value actually arrives. -
Publish/promote gate for the elicitation contract (
422 invalid_field_elicitation). Fails an unknown owner or injection phase, an applicant-owned field with neither aquestionnor adescription(the runtime would otherwise fall back to a generated phrase and ultimately to the raw field id), and a non-applicant field that does not say where its value comes from or when. Unconditional, unlike the immigration-scoped grounded-why gate: a question the applicant cannot answer is a defect in any domain. Compatibility governs silence about ownership only. A field that declares no owner below schemav2is legacy content, read as applicant-owned; fromv2it must say. Every other clause applies at every version — compatibility is a licence to omit the owner, never a licence to ship a field the runtime will have to invent a question for.
Fixed
GET /rulebooks/{id}/schemano longer discards a member field’s own authoredquestion(aethis-core#390). The overlay fell straight through todescriptionfor any field with no vocabulary entry, so a field whose applicant-facing wording had been written during authoring served its internal engineering prose instead —0.52.0corrected only the one field named in the vocabulary. Precedence is now vocabulary → compiled question → description, matching whatGET /rulesets/{id}/schemaalready did.
Added
-
The Rulebook field vocabulary is now load-bearing:
elicitation_owner+default_policyonRulebookFieldSpec, overlaid ontoGET /rulebooks/{id}/schema(tda-server#901).rb.fieldswas written byset-fields, read back by the lock handlers, and consumed by nothing — it had no effect on the field universe served to a conversational agent. Nothing incompiled_fieldssaid who may supply a fact, so every field looked equally askable, and a fact the applicant cannot know (the date the application is intended to be made — chosen by the supervising solicitor, not observed) was asked of them, invented, and then fed a statutory age exemption.elicitation_owner(applicant|caseworker|system|derived, defaulting toapplicant) declares who may supply a fact.default_policy(currently onlycase_open_date) names a policy the runtime materialises into an ordinary stored fact at a known moment — deliberately not evaluated at decision time, which would make the same case decide differently tomorrow, break decision replay (decision_id/content_digest) and violate DS-41. The overlay is a left join: a vocabulary entry enriches a compiled field and never invents one; a compiled field with no entry passes through unchanged, so a partial vocabulary is safe. A rulebook withfields: []— the deployed shape — is a behavioural no-op. Both keys are always emitted, so a consumer gets an explicitapplicantrather than having to infer askability from absence. Additive on the wire: consumers predating the keys ignore them. -
include_optimal_pathonPOST /decide, defaulting totrue(aethis-core#374, step 1 of 3).optimal_pathwas computed unconditionally on everyundetermineddecision — a constraint-solver Optimize pass measured at ~4.4s of an ~11.4s evaluation on the realaethis/form-anrulebook, roughly 38% of a conversational turn, on a fallback that consumers read only whennext_questionis absent (so on the common path it is computed, serialised and discarded). It is now behind a request flag mirroringinclude_explanation/include_trace/include_graph_overlay, and joins the decision cache key exactly asinclude_composed_explanationdoes, so a flag-off entry can never be served to a flag-on caller. The default istrue: this release is a behavioural no-op. Consumers opt out (step 2) and the default flips (step 3) as separate, later changes.
Performance
-
Version-pinned navigator cache entries no longer expire every 5 minutes
(aethis-core#374).
NAVIGATOR_CACHE_TTLapplied a 5-minute expiry to an artifact the cache’s own contract calls immutable: a pinned rulebook composition folds every memberruleset_idand the outcome-logic hash into its key, so a re-publish mints a new key rather than changing what an existing one means — the TTL could never prevent a stale read, only force a ~0.43s solver recompilation of unchanged content. Production measured 83 rebuilds against 106 cache hits over 12h (a 44% miss rate). Pinned entries now takeNAVIGATOR_CACHE_TTL_PINNED(24h) via the same per-entry TTL mechanism the decision cache already uses, and remain bounded byNAVIGATOR_CACHE_MAXeviction. Non-version-stable keys are unchanged. -
Composed-explanation build 2.4x faster; no behaviour change (aethis-core#371).
_check_condition_status— the constraint-solver core shared by the layered explanation, the graph overlay, andevaluate_subexpr_status— accumulatedstr(constraint)for every constraint it asserted, while the only consumer waslen(...)in two diagnostic messages.str()on a solver boolean expression runs the solver’s pretty-printer over the entire AST, so every rendered string was built and discarded. Now a counter. Profiled offline against the real publishedaethis/form-anmember bundles: composed explanation build 10.71s -> 4.56s, with 6.04s of the original attributable to the printer.
Added (P2)
-
include_composed_explanationrequest flag on/decide(#360). Opt-in composed (rulebook) explanation: when true on arulebook_idrequest,explanationcarries the layered per-criterion explanation across every member ruleset, with groups scoped<section>.<group>and per-criteriontitle/status/ labelledsupporting_facts/source_refs(plus per-member publish-resolvedsource_references, digest-bound to the exact member cut composed). Default false is load-bearing (spec Decision 7): downstream consumers auto-prefer a non-None explanation, so a default-on rulebook explanation would silently rewrite their user-visible narration. Composed mode omitsdecision_path— an all-sections composition has no single winning path, and naming one would be fabricated reasoning. The flag participates in the decision-cache key, so a warm entry written under one flag value can never serve the other. No-op onruleset_idrequests. -
Labelled supporting facts (#360). Explanation supporting facts are now
{field, label, value, display_value}—labeljoins the ruleset’scompiled_fields[].label(“Age at application”), anddisplay_valueis a deterministic human rendering (booleans as Yes/No, DATE ordinals as ISO dates, enum constants de-underscored) so renderers never need raw field ids. Additive:fieldandvalueare unchanged. -
failure_reasonsonnot_satisfiedcriteria (#360). “Why ruled out” is first-class: each not-satisfied criterion carries the solver’s unsat-core reasons — the raw form (expression, plus a newfieldslist naming the input fields a condition references) AND a deterministichumanstring built with the same label join, free of raw field ids and raw solver syntax. Same-deploy determinism only; cross-version unsat-core stability is explicitly not claimed (the solver returns a core, not a canonical one).
Fixed (P2)
- Failure-reason field/value mis-parse (#360). Unsat-core answer
assertions were parsed back with
rsplit("_", 1), fabricating non-existent field names for any underscore-bearing value (fieldeng.selt_provideransweredielts_ukviwas reported as fieldeng.selt_provider_ielts, valueukvi). Assertion meanings are now recorded at track time and looked up, never string-parsed;valuenow carries the actual typed answer rather than its string form.
Added
-
content_identityon the/decideenvelope (#359). One content-identity anchor for both request kinds: the leaf ruleset’scontent_digestfor a ruleset decide, or the version-pinned composition key for a rulebook decide (previously computed internally for the decision-cache key and never surfaced, so a rulebook decision had no exposed content identity at all). Replay-critical: a caller that retains it can later verify an explanation was recomputed against the same rule artefact, and refuse it if that artefact has since been superseded. The two kinds bind at different strengths. A leafcontent_digestis a true content identity (a hash of the compiled bundle). A rulebook composition key is a composition identity: it binds version-pinned member ids plus the outcome logic, not the members’ compiled bytes. Under the converged promotion model each cut pins a new bundle id, so the two coincide in practice — but a legacy path mutating content under a stable bundle id would leave the value unchanged across that change. Content-level binding for compositions is #39; until it lands, a rulebookcontent_identityis composition identity. Null is meaningful. A rulebook whose members are not all version-pinned has no stable composition identity, so the field isnullrather than a fabricated value — such a decision is not verifiably replayable, and a consumer verifying provenance must refuse it instead of readingnull == nullas a match. The value rides the cached payload like every other artefact-identity field; because the decision-cache key embedscontent_identity, a warm hit can only be an entry written under the same identity, so a cache hit reports the same anchor as the cold call that populated it. -
section_results[].ruleset_id(#359). Rulebook per-section diagnostics now name the member ruleset that produced each section, resolved from the same composition the navigator compiles. The field was declared but never populated. A section prefix with no matching member staysnull— an honest gap, never a guessed id. -
Artefact citation targets (#350).
source_targetsentries now take exactly one ofurlorartefact_source_id(both/neither → 422). An artefact target resolves against the caller’s uploaded project source (ProjectSource+ProjectSourceBlob, scoped by tenant + project) with zero network calls: row/blob/re-hash digest and size agreement is verified, the verbatim quote is checked against the retained bytes (PDF page-aware), and the emitted reference is schema v2 —target_kind: "artefact",artefact_project_id,artefact_source_id,url= the authenticated/rawdownload route. Distinct fail-closed reason codes:artefact_not_found,artefact_cross_scope,legacy_source_no_artefact,artefact_integrity_mismatch,artefact_unsupported_type(plus the existingquote_not_found/digest_mismatch). -
Snapshot-on-fetch for URL citations (#350). The exact bytes fetched for
a URL target are retained in the new content-addressed
citation_snapshotscollection (deduplicated bysha256:digest, 10 MB fetch cap unchanged); newly emitted v1 references carry the additive optionalsnapshotmarker (the snapshot digest). Previously stored v1 references are served byte-for-byte unchanged — absent additive fields are omitted, never emitted as null. -
Artefact public-exposure guard (#350). Reason code
artefact_reference_public_visibility_forbiddenon every lane to public: public-bound publish with artefact targets, ruleset visibility flip when any stored cut carries a v2 reference, rulebook visibility flip when any member does, and member attach / promote-to-live into an already-public rulebook. Private uploaded sources can never become anonymously resolvable. -
Released-consumer parse-compat fixture (#350).
tests/contract/fixtures/explain_v2_artefact_response.jsonis generated through the engine’s own response models, and CI proves the released PyPIaethis-cliandaethis-sdkparse/render a v2-bearing response without failure.
Fixed
- Every
/generatecall failed on 0.49.1 (#356).GenerationInputsgrew a sixth field (source_document_ids, part of the #331 provenance wave) while_run_generation_inlinestill unpacked five positionally, so authoring generation raisedValueError: too many values to unpack (expected 5)immediately on staging and production. Now consumed by attribute; the same function already readinputs.source_document_idselsewhere, so the unpack was the single stale site. Caught by a live canary, not by CI — a positional unpack of a NamedTuple is type-checkable but was not covered by a test that ran the inline generation path.
Released as 0.49.1.Public developer release (epic aethis-workspace#643) — production safety boundary and truthful contracts. This is the version the public docs at docs.aethis.ai describe. Until it is live,0.49.0was bumped in the repo but never tagged — the deploy script bumped again at tag time, so nov0.49.0exists. These notes describe the 23 commits inv0.46.3..v0.49.1.
api.aethis.ai serves 0.46.3 and the differences are listed on
Deployed contract status.Minor rather than patch because three documented behaviours change for
existing callers — see Changed below. Pre-1.0, a minor bump carries
breaking changes.Added
- Immutable published ruleset versions (#335).
PublishedRulesetVersiondocuments are write-once (uniqueidempotency_token, unique(ruleset_id, content_digest)backstop), fronted by a CAS-protected publication head with a generation counter and an append-onlyPublicationEventstream (publish / head-move / retire / refs-enriched / corrupt-record / quarantine). Republishing mints a new version and moves the head; it never mutates a prior version, so a version id remains replayable forever. One canonical version primitive — both the legacy publish path andpromote_ruleset_to_livego throughcut_published_version. - Real resolved identity on every leaf response (#335).
/decide,/schemaand/explainnow return the resolved immutableruleset_id(never the caller’s slug), the publishedruleset_version, and acontent_digest(sha256:<hex>). The digest is folded into the navigator and decision cache keys, so a republish under a stable slug cannot serve prior content from any cache tier — including from a Cloud Run replica that never saw the publisher’s invalidation. - Publish-validated
source_references[]on criteria (#335). A public citation is a self-locating reference validated at publish time (content digest plus verbatim-quote occurrence), not an opaque string. - Documented response truth table (#335).
aethis_core/public/contracts/decision_truth_table.pyis the canonical machine-readable table, mirrored indocs/public-decision-truth-table.mdand drift-guarded by tests. - Route- and method-scoped anonymous CORS (#333).
RouteScopedCORSMiddlewaregrants any origin (no credentials) on the documented anonymous evaluate/read routes only; everything else is restricted to trusted first-party origins, with localhost off in production. X-Aethis-Records-Omittedon catalogue reads (#335), so a truncated listing announces itself instead of looking complete.- Content-addressed source artefact retention (#331), with cited-source deletion guarded so a citation can never dangle.
Changed
Three changes are visible to existing callers:- A positive verdict can no longer coexist with blocking input errors
(#335). Non-empty
field_errorsforcesdecision: "undetermined"as the last word before the response is built, and clears misleading guidance. Per-criterion facts (explanation.groups[].status,graph_overlay.nodes[].overlay.status) still report what the engine established from the valid subset of inputs, so a criterion may readsatisfiedinside a forced-undeterminedresponse.decisionis the only outcome field; never re-derive an outcome by aggregating group statuses. - An undefined top-level request key is now rejected (#335).
DecideRequestisextra="forbid", so a misspelled or unsupported key returns a structured422 extra_forbiddeninstead of being silently ignored with a200. - Empty-extraction source uploads fail loud (#327). An oversized
(>500-page), image-only or undecodable upload returns
422and stores nothing, instead of storing an empty source withextraction_errorset and returning201— which let generation run ungrounded against a source the author believed was loaded. A partial extraction that still yields text is unaffected: stored, withextraction_errorsurfaced.
Security
-
Closed environment predicate (#333).
is_production/is_deployednormaliseprodandproduction; an unsupported value refuses to boot, andDISABLE_AUTH=truerefuses to construct on any deployed environment. The deploy pipeline sets_ENVIRONMENT=prod, which the previousenvironment == "production"comparison never matched — so the production service was silently running without theDISABLE_AUTHguard or HSTS, and admitting localhost CORS. - Fail-closed auth (#333). The all-scope mock principal is minted only when auth is disabled and the environment is non-deployed.
-
Anonymous throttling keyed off a trusted proxy hop (#333). Client
identity comes from the rightmost trusted-appended
X-Forwarded-Forentry, so a forged header rotation can no longer evade a 429. Adds per-IP and global anonymous budgets, an anonymous concurrency cap, costed expensive options, and a pre-parse body cap (413 before JSON parsing). -
Anonymous decision logging is metadata-only (#333). Hashes and aggregate
metadata only — never raw
field_valuesorcaller_ref— through a single choke point with a seeded-canary test. -
internal_onlyhonoured on the authenticated cross-tenant path (#333).visible_tonow requiresvisibility=public AND internal_only != true, so aninternal_onlyruleset wrongly flipped public is not readable by any API key on a public engine. A loud boot-time invariant scan backs it. -
No real tenant identifier in the public OpenAPI document (#337). The
caller_refexample named a real pilot firm; it is now a neutral placeholder, guarded by a test that asserts on the serialized document. -
A rulebook can no longer compose another tenant’s private ruleset
(#334).
ruleset_refswere stored verbatim from the request body, and every path that later resolved a member — the navigator (so the decision itself), the/decidegraph overlay, and the rulebook schema / fields / graph / activate reads — looked it up with no tenant or visibility predicate. A tenant could therefore point its own rulebook at another tenant’s private compiled-ruleset id (or collide on a non-namespacedsection_id), call/decideon its own rulebook, and receive a decision computed against the other tenant’s rules plus their criteria text. Closed at two independent gates: refs are validated againstvisible_toon create and update (422ruleset_not_visible), and all six resolution sites are scoped to the rulebook owner’s visibility — including a re-assertion on the compiled-bundle cache-hit path, since that cache is keyed by the compiled-ruleset id with no tenant component and would otherwise launder an unauthorised read. Composing a publicaethis/*ruleset into a private rulebook is unaffected and covered by tests.Tier-2 tenancy — this must carry an independent adversarial review before the tag is cut. It was authored and self-tested in one session, which
adversarial-review-disciplinedoes not accept as sufficient for a tenancy change. Do not treat its presence in these notes as review having happened.
Changed (continued)
- Anonymous rate limiting moved to layered rolling windows (#343). The
anonymous lane had been left on the pre-#552 shape — one counter per calendar
day, and
Retry-After: 86400on every breach — because epic #552 re-sliced the authenticated path and explicitly declined to re-tune anonymous. It now meters three rolling windows per class: burst (the current minute), sustained (rolling 1h) and daily (rolling 24h), read from one minute-bucket granularity in a single aggregation. The tightest breach is the one reported andRetry-Afternames the shortest window that actually clears, so a breach costs at most one window rather than lasting until UTC midnight.read(schema/list, near-free) now out-budgetsdecide(the compute), where both were previously capped at 500/day. The shared global cap gets the same treatment — it could previously 429 every anonymous caller until midnight. Numbers are owner-tunable policy data; this is a loosening of an abuse control and may warrant aPILOT_RELAXATIONSentry.
Fixed
- A rulebook member’s source references now survive publish → promote
(#336). Publish-time source resolution was gated on
rulebook is None, so a ruleset published as a rulebook member never resolved its declared citation keys through the publish path at all, and the cut that records them sat in the same branch. Members now resolve like any other leaf and cut a version carrying the references, deduping with the promote-time cut via the existing idempotency contract. The D2 traceability claim foraethis/uk-fsm/child-eligibilityremains unverified until a live publish → promote →/explainrun confirms it end to end. - Upload success-path test coverage restored (#339) — the invariant that the empty-extraction guard does not fire on a partial decode had no working coverage, because the only test exercising it never ran.
generate-and-testtimeout raised to a tunable 270s below the Cloud Run ceiling (#326).DSL_TEMPLATEoperator list derived from theOperatorenum, with drift-guard tests (#328).
Deploy notes
Blocking, before the tag: #334 needs an independent adversarial review.Resolved — the review ran and found real holes. An independent fresh-context reviewer returned BLOCKING on two counts: (a) the fix had no regression coverage at its highest-value seam — deleting one line (tenant_id=rulebook.tenant_id) reopened the vulnerability on/decidebyte-for-byte with the entire suite green, as did deleting the cache-hit probe call and de-scoping the probe’s query; (b) the “all six member-resolution sites are scoped” claim was false — a seventh site (_resolve_section_names) leaked another tenant’s private ruleset name to the anonymous rulebook catalogue. Both closed in #348, which also fixed a 24h cache-revocation window, a falsy-vs-is Noneinconsistency across five sites, and two latent traps. The three surviving mutations are now each caught by a named test. Shipped inv0.49.1. Independence obtained:different-session+ adversarial framing; notdifferent-provider, whichadversarial-review-disciplineasks for on Tier-2 — a second review from another model family would still be additive.- Unverified, before the claim: #336 restores the mechanism but the D2
source-traceability claim on
aethis/uk-fsm/child-eligibilityneeds a live internal-key publish → promote →/explainrun. Until it passes, that showcase is excluded from the traceability claim or the claim is narrowed. - Owner-unreviewed policy: #343’s anonymous limits were chosen without a sign-off on the numbers. They are policy data; re-tune in one line if the exposure is wrong for launch.
- Every acceptance item requiring a live revision —
check-public-deploy-security.pyagainst production with--burst/--concurrency-probe/ canary, plus the TTL-index, backup-retention and raw-canary-absence attestations — is deferred to the approved production deploy and consumed by epic #643 P10. The script emits these aspending-attestationitems with the exact canary token to search for. - After the tag is live, refresh the canonical demo bundles
(
aethis-examples:make rebuild-canonical-bundlesthenmake snapshot-canonical-bundles) — this release changes decision-envelope shape. - Update
mintlify-docs/api-reference/openapi.jsoninfo.versionand deletereference/deployed-contract.mdx(plus the caveat snippet that links to it) onceengine_versionreports0.49.0.
Publish with citations — and be able to cite a document you hold, not just one
the internet happens to host.
-
feat:
aethis publish --source-targets <file>resolves a ruleset’s citation keys. A YAML or JSON targets file maps each citation key your criteria declare (source_refs) to the document it cites: title, authority, licence, and the verbatim quoted text. The engine verifies every quote against the source bytes at publish time and rejects the publish if any citation fails — there is no half-cited ruleset. -
feat: a citation can point at a file you uploaded, not only a public URL. An entry naming
file:is uploaded to the project and cited by itssource_id, so the rules can cite the very documents they were generated from. The engine resolves it from retained bytes with no network call at all. -
feat: an identical file is never uploaded twice. File targets are matched by
sha256against the project’s existing sources and reused when the bytes are already there — including two entries in the same run naming byte-identical files. The API does not deduplicate uploads, so re-running a publish previously grew the project a duplicate source per citation. -
feat: a malformed targets file costs no round trip. Exactly-one-of
url/file, a readable file, anhttps://URL, the requiredtitle/authority/licence/quote.exact, and unknown fields are all checked locally — every problem in the file reported at once, before the first API call, so nothing is uploaded against a targets file that was never going to publish. -
feat: an uploaded-artefact citation is never rendered as though it were a public link.
aethis decide --explainandaethis explainlabel these references as an uploaded snapshot verified at publish, state that the download is authenticated and needs a key withprojects:readon the project, and resolve the engine-relative download path against the host you called — while a URL citation keeps reading as the public link it is.aethis publishreports the same distinction per target as it resolves them. -
feat:
aethiscan list a project’s uploaded sources (AethisClient.list_sources), which is what makes the digest comparison above possible. -
fix: a rejected citation now says which one and why. Publish-time citation resolution is fail-closed and the API itemises every failing key (
{source_id, reason_code, message}), but the CLI collapsed the whole envelope to its summary line — an author with three citations and one wrong quote learned neither which key failed nor what was wrong with it. Every itemised failure is now printed under the error. Applies to any endpoint returning afailureslist, not just publish. -
feat: citation targets that never landed are reported, not silently dropped. The engine only resolves the citation keys the compiled ruleset actually declares, and ignores the rest — so a mistyped key published “successfully” with zero citations attached, after uploading the files. After a publish with
--source-targets, the CLI reads the published ruleset back and says how many targets landed, naming any that did not. A failed read-back degrades to an honest note; it never turns a successful publish into a failure. - feat: a failed publish says your uploads are still there. Publishing is fail-closed but the uploads that preceded it are not rolled back, and re-running reuses them by digest rather than duplicating them. The error now says so, instead of leaving the state of the project a guess.
- fix: a duplicate citation key is rejected instead of silently overwriting. YAML and JSON both let the last definition of a repeated key win; in a citation manifest that quietly discards a document and publishes the other one into an immutable ruleset. Both formats now refuse duplicate keys (at any depth in the file).
-
ci(publish): the downstream-unstick sweep now reaches every consumer repo. The
unstick-downstreamjob searched a single owner, so a PR carrying anaethis-needs: aethis-climarker in a repo under a different owner was never found and sat as a draft indefinitely after the release it was waiting for went out. It now queries each consumer repo individually withgh pr list --repo, which is genuinely repo-scoped, and matches the marker against the PR body returned by that same call (one request per repo instead of a search plus a fetch per hit). A repo the token cannot read is reported as a warning instead of failing the whole sweep. No package or runtime change.
Makes the SDK a safe, immutable release component for the public developer
release (epic aethis-workspace#643, P9 / aethis-sdk-python#29). Three classes of
“looks fine, isn’t” are closed at the type level, and the release itself now
carries verifiable integrity evidence.
Replay identity: absence no longer looks like a version
- breaking (behavioural):
DecideResponse.ruleset_versionisstr | Noneand no longer defaults to"unknown". The engine reports an unresolved version as the literal string"unknown"(a rulebook call, or an artefact published before immutable versions); the SDK also defaulted to that string, so a caller writing an audit record got a plausible-looking version whether or not anything had been resolved. Every unresolved sentinel ("unknown","","none","null","n/a", case-insensitive) now normalises toNoneonruleset_version,content_digest,ruleset_id,engine_version,decision_idandinputs_hash. Code readingresponse.ruleset_versiongetsNonewhere it previously got"unknown". - feat:
content_digestonDecideResponse, andruleset_version+content_digestonSchemaResponse. The resolved immutable identity aethis-core stamps on/decide,/schemaand/explain(aethis-core#330). - feat:
ReplayIdentity/ContentIdentity+require_replay_identity()/require_content_identity(). These return a complete identity or raiseAethisReplayIdentityErrornaming exactly which parts are unresolved — so recording an incomplete audit reference is an explicit act, not a default. The soft accessorsreplay_identity/content_identityreturnNoneinstead of raising.
Blocking errors cannot become a completed or positive result
- feat:
DecideResponse.blocking_errors(always a mapping),.has_blocking_errors,.is_terminal,.raise_for_blocking_errors(). The engine suppressesnext_questionwhile blockingfield_errorsare outstanding, so a blocked response is byte-shaped like a finished one on that field.is_terminalis the honest check. - feat: the parse boundary refuses a self-contradicting envelope. A 2xx reporting
eligible/not_eligiblebeside non-emptyfield_errors— or an embedded copy inexplanation.decision/trace.statusthat contradicts the headline — raises the newAethisContractViolationrather than becoming an object a caller acts on. Enforced in the model, so it holds on the sync client, the async client and the sessions alike. - feat:
SessionStatusgainsfield_errors,replay_identity,.blocked,.is_completeand.raise_if_blocked(), with a constructor invariant that makes a positive-and-blocked status unconstructible. Sessions gainblocking_errors()andis_complete()(sync and async). - feat: new
AethisFieldErrorsexception carrying.field_errors, raised by the opt-inraise_for_blocking_errors()/raise_if_blocked()guards.
Typed source provenance
- feat:
SourceReference+SourceQuotemodels — the publish-validated citation contract (source_id,title,authority, HTTPSurl,locator,source_version,source_date,content_digest,licence,verified_at, verbatimquote, self-locatingdeep_link,schema_version), returned identically by both explanation surfaces. Unknown fields are preserved so the additiveschema_versionevolution cannot break a pinned consumer. - feat:
get_explanation(ruleset_id)(sync + async) returning the typedExplainResponse, with resolved identity and typed references per criterion.explain()still returns the raw dict for existing callers. - feat:
DecideResponse.decision_explanation+.source_references()parse the/decideexplanation intoDecisionExplanation. Note the two surfaces differ:/explainreturns a flatcriteriaarray,/decidenests criteria underexplanation.groups[].criteria[]. They share the DTO, not the envelope — the SDK models them separately and the tests assert the distinction.
The two access boundaries are labelled
- feat:
AethisError.boundaryis"evaluation"or"authoring"on a 401/403, and the exception message now names which door was closed — no-key evaluation (/decide,/rulesets,/schema,/explain) versus invite-only authoring — plus the access-request URL. README and examples carry the same labelling before either path.
Release integrity and hermetic install evidence
- feat:
scripts/release_integrity.pyemits the tuple a release candidate is pinned on —(package, version)→ exact sdist/wheel sha256 → source commit/branch/clean-state — and re-verifies it against local files or against what PyPI actually serves. Wired intopublish.ymlbefore publication (--require-clean) and after (--verify-registry). - feat:
scripts/hermetic_install_check.pyinstalls the exact artefact into a throwaway world — temporaryHOME/XDG_*/cache, everyAETHIS_*and provider key unset, empty cache on first install, no alternate index — then runs an offline smoke that parses captured engine payloads through the installed package, and a poisoned-artefact negative control that must fail. NewhermeticCI job across ubuntu/macos × Python 3.11/3.12/3.13. - feat:
scripts/capture_engine_fixtures.pyrecords the fixtures undertests/fixtures/from a live engine (anonymously, against a public showcase ruleset), including the engine’s own JSON Schemas, so the mocked suite is tested against real wire payloads rather than hand-written approximations. - chore: Python 3.13 added to the classifiers and the CI matrix.
jsonschemaadded to thedevextra (test-only; the shipped package is still justhttpx+pydantic).
Review-wave fixes (same release)
- fix:
get_source()reported the wrong access boundary./rulesets/{id}/sourcesits under the/public/rulesetsprefix but is key-required behind a scope external keys are not issued, so the prefix match labelled a 401 there"evaluation"and told the reader to go looking for a ruleset-visibility problem for a door that will never open. Key-required sub-paths are now excluded from the evaluation prefix, and bothget_sourcedocstrings say so. Preferget_explanation(), which is anonymous on a public ruleset and returns the sameSourceReferenceDTO. - fix:
content_digestis validated against^sha256:[0-9a-f]{64}$.md5:…, a truncatedsha256:beef, non-hex, uppercase and bare-hex values now normalise toNonerather than being carried into an audit record — the same “looks like a value, isn’t” class as the"unknown"version. - fix:
--require-cleanpassed vacuously when git provenance was unreadable.source_provenance()returnsdirty: None(unknown) when git cannot be read, and the gate tested it for falsiness — so with no.gitthe script exited 0 while recordingcommit: null. Provenance problems are now checked positively (dirty is not False, plus a missing commit in its own right). - fix: the poisoned-artefact control was vacuous on every CI runner. It flipped the final byte, which lands in the end-of-central-directory comment field; strict zip readers reject it, lenient ones scan backwards and install happily. It is replaced by two controls: a digest control (a valid, installable substituted wheel that the real
verify_filesmust reject — deterministic, and the layer that actually protects users) and an installer control (a wheel whose compressed stream is corrupted, asserted positively onstage == "uv pip install"plus stderr that names the corruption, rather than on the absence of one unrelated string). - fix: publication is gated on
--verify-filesimmediately before the upload, which is irreversible, in addition to the post-publish registry check. - fix: the README session loop no longer demonstrates the bug this release exists to prevent. It looped on
next_question() is not None; it now loops onstatus(), with a table of the four states that loop conflated. - fix: ruff’s rule selection is pinned (
ruff>=0.6.0,<0.17plus an explicit[tool.ruff.lint] select). CI installed ruff unpinned, so 0.16.0’s wider default turned the lint gate red on 61 errors — before pytest ran at all. Mirrors aethis-cli#90. CI also now lintsexamples/, which it had never covered. - fix: captured OpenAPI
examplesare stripped, and a test guards the fixtures. The engine’scaller_refexample named a real pilot firm, which a verbatim capture committed to this public repo.examplescarry no structural information and validation ignores them, so they are no longer captured; a standing test fails on any internal tenant name or immigration term intests/fixtures/. - fix: the wheel is now byte-reproducible. Three builds of the same clean tree produced three different digest pairs, so the tuple’s “which commit produced these bytes” leg was an attestation nobody could re-derive.
SOURCE_DATE_EPOCHis now set from the commit timestamp in both build workflows, a CI step rebuilds the wheel and fails if the digest moves, and the tuple records precisely what is verifiable — the sdist is still not reproducible (setuptools varies the archive), and says so rather than implying otherwise. - test: the two invariants that no ordinary test could reach are now covered.
is_terminal’s blocking-error clause andSessionStatus.is_complete’snot blockedclause are each masked by the parse validator and the constructor invariant respectively — so deleting either left the suite green while removing the last guard on a bypass route. Both are now exercised throughmodel_construct/object.__setattr__/dataclasses.replace, and verified to fail when the clause is removed.
Safety and provenance for everything the CLI reads back from the API.Versioning note. This release changes two documented behaviours —
aethis explain --output json emits the whole envelope rather than the bare
criteria array, and a blocked evaluation now exits 3 where it previously
exited 0. Under strict SemVer a breaking change is a major bump; taken as a
minor here because the package is pre-1.0 (0.x), where the published
rule is that minor carries breaking changes. Recording the call explicitly
rather than leaving it to be inferred: both changes replace behaviour that was
unsafe (a script could not tell a rejected input from a decision), which is
why they ship rather than waiting for 1.0.-
feat: a rejected input can never look like a result. When a decision response carries blocking
field_errors,aethis decideandaethis rulebooks decideprint the rejected inputs instead of a verdict and exit3(new exit code:0decided,1call failed,3inputs rejected). JSON output reports"decision": "undetermined"and records the block underaethis_cli_contract. The CLI enforces this rather than trusting it: if a server ever returnseligiblebeside blocking errors — a stale deployment, a caching proxy, a third-party API-compatible server — the contradiction is overridden, reported, and never rendered as success in human output, in JSON, or through the exit status. Newaethis_cli.contractmodule owns the rule. -
feat: immutable identity on every decision surface. Human output gains a
Ruleset identityblock (ruleset id, published version,sha256:content digest, engine version, decision id, inputs hash) so a decision can be reproduced or audited later. Aruleset_versionofunknown— which a published ruleset must never report — is called out as unresolved instead of printed as though it were an identity. -
feat: supporting sources are shown, and kept separate from the rules. Where a ruleset publishes validated source references,
aethis decide --explainandaethis explainrender them under their ownSourcesheading: document title, authority, locator, the verbatim quoted text, deep link, licence, verification time and source digest. A reference that arrived incomplete is flagged rather than rendered as a confident-looking citation. When a ruleset publishes none, the output says so rather than showing nothing. - feat: output distinguishes result, logic trace and source. The three now have their own headings, and the logic trace is labelled as explanatory — per-criterion statuses answer “what is true of this criterion”, never “what may I act on”.
-
change:
aethis explain --output jsonnow emits the whole envelope (ruleset_id,slug,ruleset_version,content_digest,criteria) instead of the barecriteriaarray. Provenance on the machine-readable path is the whole point of the identity contract; a script that consumed the old shape needs.criteria. -
feat: undeclared
/deciderequest options are refused locally. The API rejects an unknown top-level request member with a 422 rather than ignoring it; the CLI now names the offending option before spending a round-trip, and renders the server’s validation envelope readably if one is ever returned. -
feat: the capability boundary is visible before you hit it. Root help,
decide/explainhelp, the README and every auth-required error now state plainly that evaluation needs no account and no key, and that authoring is invite-only (with the access link). -
feat: release integrity and hermetic first-install evidence.
scripts/release-integrity.pybinds the exact published bytes to the commit they were built from (version + sdist/wheelsha256+ source commit) and can re-check that against the files the registry serves.scripts/hermetic-install-check.pyproves the CLI works for someone who has never run it: temporary HOME/XDG/config/cache, no Aethis or provider credentials, empty-cache install from one source only, across supported runtime/OS/architecture — with a poisoned-cache negative control that must fail. Both run in CI on every PR (Linux + macOS, Python 3.11/3.12/3.13) and on every release, before and after publication. -
feat: the guard matches the engine’s own forcing sweep exactly. When a response is blocked, all five embedded copies of a terminal verdict are scrubbed — top-level
decision,explanation.decision,explanation.decision_path,trace.status,trace.path— so a blocked result can never be printed above a green “Satisfied by: …”. A--json <fields>projection can no longer drop theaethis_cli_contractrecord either. -
fix: a non-JSON response body is an error, not a traceback. A 2xx carrying HTML (an intermediary’s error page, a truncated body) now surfaces as one readable API error instead of a
JSONDecodeErrorescaping mid-command. -
fix: documented commands that would not run.
--outputis a root option and must precede the subcommand; five documented invocations had it after (including the README’s flagship shell-gate example, whoseelsebranch therefore misreported). All fixed, and a test now resolves every documented invocation against the real command tree so this class cannot come back. -
test: the contract oracle is itself gated.
scripts/mutation-check.pybreaks the contract sixteen different ways — deleting each scrub site, redefining the blocking exit code, making the blocking predicate always-false, unguarding the rulebook surface — and requires the suite to go red for every one. It runs in CI. -
test: contract fixtures are captured, never hand-written. The new tests run against payloads recorded from a live engine (terminal decisions, each class of blocking input error, an incomplete evaluation, the 422 for an undeclared request member, the explain envelope) plus source-reference DTOs serialised by the engine’s own model. Regenerate with
scripts/gen-contract-fixtures.py; provenance is documented intests/fixtures/contract/README.md. -
fix: no spurious traceback when a login callback is cancelled or times out. The OAuth callback server closed its socket while its background thread was still waiting on it, so the thread died on
ValueError: Invalid file descriptor: -1and printed an unhandled-exception traceback afteraethis logintimed out. The serving thread is now signalled before the socket closes, and treats a closed socket as its exit condition rather than an error.
- security: every tool’s server-supplied free text is fenced as untrusted
data. JSON-passthrough tools (list/discover/schema/graph/rulebook and
friends) now wrap their whole response in a single
<api_response>fence with the untrusted preface, so a server- or tenant-authored free-text field (name, description, domain, message, …) can no longer smuggle instructions to the model via the JSON blob. Prose tools keep their per-field fences. A new deterministic serializer-coverage test drives every tool with taint sentinels and fails if any free-text leaf is ever emitted outside a fence. Closes the untrusted-JSON-passthrough gap (aethis-mcp#45). - security: capability annotations + containment. Every tool now carries MCP
annotations (
readOnlyHint/destructiveHint/openWorldHint) derived from a single capability registry, so a host renders correct read/destructive hints and can gate approval on the mutating tools. A test enforces that no no-API-key (anonymous) tool has any mutation capability, and that the registry matches which handlers actually require a key. - release:
server.jsonis derived frompackage.jsonand drift-guarded.server.json(the official MCP Registry record) now tracks the package version source of truth;npm run check:server-json(run in the test suite/CI) fails on drift. Corrected the staleserver.jsonversion (0.5.1 → current). - release: generated tool inventory.
tool-inventory.jsonis a generated, drift-guarded listing of the tool surface, used to verify a fresh install and as the source of truth for the published tools reference. - release: gated, evidence-producing publish pipeline. An unprivileged build stage produces one immutable tarball with a sha256 digest, a CycloneDX SBOM and a candidate manifest; npm and the official MCP Registry are then published via separate protected environments (named reviewer), workflow-bound OIDC and no stored token. Post-publish verification and a clean-environment fresh install must both confirm the exact name/version before a release reports success. See docs/RELEASE.md.
- docs: align Simpson paper citations with v3.13 (issue #53). The construction-insurance demo now cites the current paper version (v3.13, 2026) instead of v3.11; removes any presentation of the withdrawn GPT-5.4 low-reasoning-effort 7/11 figure as a live result (the v3.8 withdrawal note remains as historical context); and removes configuration-level API detail (parameter names) from the benchmark methodology text. All real benchmark numbers are unchanged.
- feat:
usage()+client.rate_limit— rate-limit budget + headers. NewAethis.usage()/AsyncAethis.usage()return aUsageResponse(per-operation-classused/limit/remaining/resetover the rolling 24h window + a 7/30-day rolling summary) fromGET /api/v1/public/usage. Every response’sX-RateLimit-*headers are now parsed ontoclient.rate_limit(aRateLimitmodel:operation_class/limit/remaining/reset), so a consuming app can read its remaining budget — especiallygenerate(the scarce LLM class) — without a separate call. New modelsUsageResponse,ClassUsage,RollingUsage,RateLimit, all exported. (epic aethis-workspace#552) - Requires aethis-core’s
/usage+X-RateLimit-*surface (epic #552 P2); the public release of this version is held until that is live onapi.aethis.ai. - ci: cut a GitHub Release on publish. The
publishworkflow now creates a GitHub Release for each just-published tag, using that version’sCHANGELOG.mdsection as the release notes, so the “watch → releases” subscribe channel stays current automatically. Idempotent (create-or-skip on an existing Release) and--verify-tag(never mints a synthetic tag). No package/runtime change. (epic aethis-workspace#526)
- feat:
aethis_usagetool. Reports the caller’s rate-limit budget per operation class (decide / generate / author / read / keys / admin) over the rolling 24h window — used, limit, remaining, reset — fromGET /api/v1/public/usage, so an agent authoring inside Claude Code / Cursor / Windsurf can see and report the developer’s remaininggeneratebudget before a 429. NewAethisClient.usage(); tenant-scoped (requires an API key). (epic aethis-workspace#552) - Requires aethis-core with the
/api/v1/public/usageendpoint live (epic #552 P2). The public npm release of this version is held until that endpoint is live onapi.aethis.ai.
- feat:
aethis usage— show your rate-limit budget per operation class (decide / generate / author / read / keys / admin) as a table: used / limit / remaining / reset, over the rolling 24h window.generate(LLM rule generation) is the scarce class; browsing and status polling (read) are effectively unlimited.--json/piped emits the raw/usagepayload. NewAethisClient.usage(). - feat: remaining generate budget after
aethis generate. The CLI now reads theX-RateLimit-*response headers (captured on every request asAethisClient.last_rate_limit) and, after a generation is queued, prints “N generations left in the current 24h window” — so a 429 is never the first signal. The line turns yellow at ≤5 remaining. - Requires aethis-core with the
GET /api/v1/public/usageendpoint +X-RateLimit-*headers live (epic aethis-workspace#552, P2). The public release of this version is held until that surface is live onapi.aethis.ai.
Metering & rate-limit revamp (epic aethis-workspace#552) — P4 over-limit rate.
Added
- Per-key over-limit rate is now persisted and queryable (aethis-core#320
follow-on). A new
over_limit_countfield on the hourlyrate_limitsbucket is incremented whenever a request is observed over its class limit — in report-only mode too, since the observation is the oracle the Bridge tunes limits against (not just a log line). The write is best-effort and fully exception-isolated: a failure can never alter the rate-limit decision or 500 the request.GET /api/v1/admin/usage/overviewnow returnsover_limit_24hper(key, class)— the rolling-24h over-limit count — so the console usage dashboard (ok_swift#665) and the approaching-cap alert (godseye#46) can show the 429-rate metric alongside proximity-to-cap.
Metering & rate-limit revamp (epic aethis-workspace#552) — P4 admin primitive.
Added
GET /api/v1/admin/usage/overview(aethis-core#320) — internal-gated (admin:readscope ANDinternal==true, same stack as the rest of/api/v1/admin/*), read-only, cross-tenant. Returns, per(key_id, tenant_id, tier, class), the rolling-24husedjoined against the tier×class limit withremainingandpct_of_limit, ranked by proximity to cap descending (bypct_of_limit, never rawused— so a near-cap low-tier key is never lost behind a high-volume internal one) so “approaching cap” is the top of the list. Filters:tier,key_id,class,min_pct(the alert threshold),limit;total_matchingreports the pre-limitcount so a caller knows when the display was clamped. The rolling-24husedis the same hourly-bucket sum the enforcement path andGET /public/usageuse, so the overview can never disagree with per-key usage. This is the shared primitive bothok_swift#665(console usage dashboard) andgodseye#46(approaching-cap alert) consume. Stays in the live/openapi.jsonbut out of the published public contract. Over-limit history (429-rate) is a deliberate follow-on: therate_limit_over_limitevent is log-only, so a history feed needs a metering write-path change.
Metering & rate-limit revamp (epic aethis-workspace#552) — reaches production.
Added
GET /api/v1/public/usage— calling-key-scoped usage per operation class (rolling 24hused/limit/remaining/reset, plus 7-day and 30-day rolling totals). Deliberately unmetered.X-RateLimit-*response headers (Class/Limit/Remaining/Reset) on metered endpoints — forward visibility so callers can see their budget without waiting for a 429.
Changed
- Rate-limit counters re-sliced into six operation classes —
decide,generate,author,read,keys,admin(one counter is both the limit and the usage metric). Policy is expressed as tier×class data. - Meter the scarce thing: only
generate(LLM-backed) carries a real ceiling. The old sharedprojectsauthoring bucket (2000/day — the squeeze that made a large rulebook “hit the 500 limit”) is split intoauthor(20k/day) andread(100k/day), so ordinary authoring no longer competes with generation. - Rolling 24-hour window (hourly sub-buckets) replaces the calendar-day reset.
Notes
generateenforcement ships report-only (RATE_LIMIT_ENFORCEenv, default off): the scarce-class limits are logged (rate_limit_over_limitevent) but do not reject, pending owner tuning on real 429-rate evidence. Every other class is preserved-or-loosened vs 0.46.2, so no traffic that succeeds today is newly rejected.
- Startup update-check nudge. On startup, the server checks the npm
registry’s
latestversion foraethis-mcpand, if a newer release exists, writes a one-line notice to its stderr log (visible in your MCP host’s server logs) pointing at the Releases page for what’s new (workspace epic #537, aethis-mcp#61). Non-blocking — the check runs in the background and never delays server startup — and fail-silent on any network error or timeout. Opt out withAETHIS_DISABLE_UPDATE_CHECK=1(mirrors the same variable in aethis-cli); also skipped automatically whenCIis set. - CI: cut a GitHub Release on publish.
publish.ymlnow creates a GitHub Release for each published tag, using that version’s CHANGELOG section as the notes body (idempotent create-or-skip). Introduces the Releases channel on this repo — the subscribe-able “watch → releases” channel for the unified developer changelog (workspace epic #526, aethis-mcp#59). CI-only; no runtime or package change on its own.
- feat: “what’s new” on
aethis update.aethis update/aethis update --checknow shows the changelog entries between your installed version and the latest release (titles + notes, newest ≤5, long notes truncated), sourced from the project’s GitHub Releases. If the Releases API is unreachable, rate-limited, or has nothing in range, it falls back to a link to the Releases page — the command never errors or hangs on this. The exit-time update banner also gained a “what’s new →” link to the same page. Coverage is forward-fill: only releases cut from here on populate the range, so an old install may see a gap. Newupdate_check._fetch_github_releases(); the display logic lives inupdate_cmd._releases_in_range()/_print_whats_new(). (epic aethis-workspace#537) - ci: cut a GitHub Release on publish. The
publishworkflow now creates a GitHub Release for each just-published tag, using that version’sCHANGELOG.mdsection as the release notes, so the “watch → releases” subscribe channel stays current automatically. Idempotent (create-or-skip on an existing Release) and--verify-tag(never mints a synthetic tag). No package/runtime change. (epic aethis-workspace#526)
Adds the Authoring Coach surface to MCP (aethis-mcp#57, workspace epic #514) —
skill-building feedback for rule authors, advisory only, never a gate.Engine gate: the
POST /api/v1/public/projects/{id}/review endpoint and the
ambient review_hint fields are produced by aethis-core (epic phases P1/P4).
This release must not be published to npm until that endpoint is live on
api.aethis.ai; a released client calling a not-yet-deployed route would 404.aethis_review_project(new tool). Reviews an authoring project against the deterministic authoring-coach rubric and renders the report: a score, per-check evidence across grounding / process / lifecycle, strengths, and the single highest-leverage next skill. Advisory only — it never blocks publishing. The deterministic layer needs no LLM key;coach=true(with an Anthropic key, via the usualanthropic_key_env/anthropic_key_keychain/anthropic_keyforms) adds an opt-in LLM-synthesised coaching narrative on top. All server free-text (evidence / strengths / next-skill message / coaching) is fenced withfenceUntrustedbefore it reaches the model.- Ambient
review_hintrender.aethis_generate_and_test,aethis_refine, andaethis_publishnow render a one-line coach hint when the server includes one on the response. The hint is computed entirely server-side (aethis-core P4); the client only renders it (fenced), never computes it. X-Aethis-Client: mcp/<version>on every request. The client now sends a per-surface identifier header so the engine can attribute telemetry (e.g.review_hint-shown counts) to MCP vs CLI vs SDK.- 31 tools, up from 30.
tests/tool-endpoint-map.tsand the drift suite are updated in the same change. Note: the drift suite’s live-alignment checks stay red against staging until the/reviewendpoint deploys there (expected epic ordering); the offline structural checks pass. - Tests. New mocked unit coverage in
tests/client.test.ts(reviewProject request shape, the client-id header) andtests/server.test.ts(aethis_review_projectrender + fencing, coach key resolution, ambient hint render on generate/publish).
- feat:
aethis review [<project>]— the Authoring Coach report for a project. Runs the server-side rubric and prints an authoring score, 2–3 evidence-cited strengths, and the single highest-leverage next improvement (with its docs link and the lever that fixes it). Defaults to the current project in.aethis/state.json; pass aproj_…id to review any of your projects from anywhere.--verboseshows the full per-check table;--json(and any piped/--output jsoninvocation) emits the rawReviewReport. The deterministic report needs only your API key;--coachopts into LLM mentoring prose billed to your own Anthropic key (ANTHROPIC_API_KEY). Advisory only — the exit code is always 0 regardless of score. NewAethisClient.review(). - feat: every request now sends
X-Aethis-Client: cli/<version>so the server can attribute per-surface telemetry (CLI vs MCP). The header carries no credentials and no PII, and is set once at client construction for all commands. - Requires aethis-core with the
/api/v1/public/projects/{id}/reviewendpoint live (epic aethis-workspace#514, P1). The public release of this version is held until that endpoint is live onapi.aethis.ai.
- feat(models):
robot_hints+engine_versionon the rulebook schema;engine_versionon the ruleset schema. NewRulebookSchemaResponsemodel (rulebook_id,sections,fields,robot_hints,engine_version) forGET /api/v1/public/rulebooks/{id}/schema—robot_hintsis the rulebook’s natural-language conversational-agent guidance keyed by beat (general_context,preamble,session_start,postamble,session_end,stuck),Nonefor a rulebook authored before the field existed.SchemaResponse(ruleset schema) gainsengine_version: str | None = Nonefor parity, also back-compat (defaultsNonewhen the engine doesn’t send it — true of the ruleset schema route today). - feat(models):
graph/GraphResponsefor the new/graphendpoint. NewGraphResponse(ruleset_id/rulebook_id,slug,name,graph,mermaid) andRulesetGraph(nodes,edges,sections,stats) model the ruleset/rulebook dependency graph (field → criterion → group → outcome) plus its rendered Mermaid diagram. Node/edge shape varies by nodetype, so nodes/edges stay loosely-typed dicts rather than a rigid per-type schema — deliberately permissive so a legacy or empty graph (nodes: []) still parses. - feat(client):
get_graph(ruleset_id)(sync + async) — wrapsGET /api/v1/public/rulesets/{id}/graph, returningGraphResponse. Public rulesets can be inspected without an API key, same asget_schema(). - feat(decide):
include_graph_overlayparameter ondecide()/decide_rulebook()(sync + async), and a matchinggraph_overlay: dict[str, Any] | None = Nonefield onDecideResponse. Setinclude_graph_overlay=Trueto get this decision’s per-criterion status stamped onto the ruleset’s dependency graph, in the same shapeget_graph()returns. - All additions are additive and backwards-compatible: every new field defaults to
None/False/an empty collection, so a legacy response (norobot_hints, noengine_version, nograph_overlay) still deserialises unchanged.
Propagates the aethis-core 0.37–0.40 authoring batch to the MCP surface
(aethis-mcp#49, workspace epic #327). Engine gate: live on
api.aethis.ai
0.45.2, confirmed via the drift suite’s live-alignment checks before this
release.aethis_graph(new tool). Fetches the ruleset-map graph — either for a single published ruleset (ruleset_id, may be public/anonymous for a public showcase ruleset) or a composed rulebook (rulebook_id, always requires an API key) — the same mutual-exclusivity shape asaethis_decide. Returns{ruleset_id|rulebook_id, slug, name, graph: {nodes, edges, sections, stats}, mermaid}: each node’sdisplay.sentence/display.routes/display.exprshows how that branch composes, andmermaidis a ready-to-render diagram string.include_graph_overlayonaethis_decide(additive). Stamp a specific decision’s per-criterion outcome (satisfied/not_satisfied/pending) onto that same graph and return it asgraph_overlayin the decide response — a “you are here” map for those inputs. Off by default; the response is unchanged when omitted.aethis_create_rulebook/aethis_update_rulebook(new tools). Create an empty draft Rulebook (name/domain/slug/description) or update one, both acceptingrobot_hints— beat-keyed natural-language guidance for the conversational agent (active beats:general_context,preamble,session_start,postamble,session_end,stuck; reserved:persona,conversational_style,section_transition). An unknown beat is rejected client-side before the round-trip, mirroring aethis-cli’s_validate_robot_hints(v0.23.0). Rulebook composition (outcome_logic,ruleset_refs) is a larger surface not covered by these two tools yet.years_betweenin the DSL helper reference (README). Documents the new completed-whole-years, leap-correct date operator (mirrors aethis-coreOperator.YEARS_BETWEEN, commit3607558) alongsidedays_betweenso generation can use it for age-from-date-of-birth instead ofdays_between(...) / 365(division isn’t supported anyway).- 30 tools, up from 27.
tests/tool-endpoint-map.tsand the drift suite are updated in the same change; every new operation/field/param is verified against the liveapi.aethis.aiOpenAPI document (engine 0.45.2). - Tests. New mocked unit coverage in
tests/client.test.ts/tests/server.test.tsfor the graph client methods, the create/update rulebook client methods,robot_hintsbeat validation (known + unknown + reserved), andinclude_graph_overlaypass-through; the nightly staging integration lane gains a realaethis_graphfetch, aaethis_decide include_graph_overlayround-trip, and anaethis_create_rulebook→aethis_update_rulebookrobot_hintsround-trip (with best-effort archive cleanup of the probe rulebook).
- feat(rulebooks):
aethis rulebooks graph <id>— fetch and render the rulebook-level ruleset-map dependency graph (field -> criterion -> group -> outcome). Prints a node-count summary + a table of nodes (id, type, the criterion’s human-readabledisplay.sentence, field count);--mermaidprints the raw Mermaid diagram source for piping into a renderer;--output jsonreturns the full payload ({rulebook_id, graph: {nodes, edges, sections, stats}, mermaid}), including each node’sdisplay.routes/display.exprfor programmatic consumers. This endpoint requires a valid API key even for a public rulebook (confirmed against the live engine) — unlike the ruleset-level graph below, there’s no anonymous path. NewAethisClient.get_rulebook_graph(). - feat(rulesets):
aethis rulesets graph <ruleset_id>— the single-ruleset analogue, open for public rulesets with no API key required (load_client_or_anon). Same table/--mermaid/--output jsonshape. NewAethisClient.get_ruleset_graph(). - feat:
--include-graph-overlayonaethis decideandaethis rulebooks decide— stamps the decision’s per-criterion status onto the rule-map graph, returned as agraph_overlayfield on the response (--output jsonto inspect it). Additive request flag; a plain-text hint is printed when the overlay is present and JSON wasn’t explicitly requested. - feat(rulebooks):
aethis rulebooks schemasurfacesengine_version. The schema response already carries the aethis-core build that served it (e.g.aethis-core@0.45.2); the CLI now prints it as a header line ahead of the schema payload instead of leaving it buried in the JSON. - Requires aethis-core 0.40.0+ (live on
api.aethis.ai/staging.api.aethis.aias of this release) for/graph,include_graph_overlay, andengine_versionon/schema.robot_hints(shipped v0.23.0) is unaffected by this release.
Fixed
- An API key with
expires_atset no longer 500s every request: the Mongo-stored (tz-naive) expiry is normalized to UTC before comparison, so an expired key gets its intended 401api_key_expired. Latent since the expiry field existed — no key carried a non-None expiry until 2026-07-16 (issue #275).
- feat(errors): typed 401/403/429 exceptions carrying the structured error envelope.
classify_responsenow raisesAethisAuthError(401),AethisPermissionError(403), orAethisRateLimitError(429) — each a subclass ofAethisAPIError, so existingexcept AethisAPIErrorhandlers keep catching them (non-breaking).AethisErrorgains.reason_code,.missing_permissions, and.hint, lifted out of the public API’s structured envelope ({"detail": {"error", "reason_code", "missing_permissions", "hint", ...}}), so a caller can branch onerr.reason_code == "denied_missing_permission"or readerr.missing_permissionswithout re-parsingerr.body. Plain-string and FastAPI-422-list details are untouched (fields stayNone/[]). Constructor stays backwards-compatible (new args default toNone). - test(staging): live integration lane against
staging.api.aethis.ai. Newtests/integration/(markerstaging, excluded from the PR gate) mints a real API key the way a user does — Clerk sign-in ticket → frontend-API JWT →POST /api/v1/keys/→ teardown — and exercises every public method onAethis+AsyncAethis(decide, decide_rulebook, list_rulesets, get_schema, whoami, explain, explain_failure, get_source, sync/async session flows) plus live 401/403 typed-error assertions and a contract cross-check. Reports red (never green-by-skip) when creds are missing or staging/contract is unreachable. - test(parity): recorded-live fixture parity.
tests/shapes.compare_shapediffs the mockedconftestfixture builders (make_decide_response,make_schema_response,make_ruleset_summary) against real staging payloads so the mocked suite can’t silently drift from reality; the builders were updated to match the current engine shape (slug/rulebook_id,graph_overlay/timing, richernext_question/schema fields). - chore(ci): coverage floor (
--cov-fail-under=45) +stagingmarker + nightlystaging-integration.yml(report-only,workflow_dispatch+schedule, uploads aqa-run-recordartifact for thesdk-staginglane). The coverage flags live in the CI command, not inaddopts, so a barepytest/uv run pytestworks withoutpytest-cov(which is only in thedevextra) installed; the floor is still enforced in CI.
Test-infra only — no runtime/behaviour change to the server or its tools.
- Tool-schema drift suite (
tests/drift.test.ts). Guards that the 27server.tool()input schemas never silently drift from the engine. Reads each tool’s real zod shape (no vendored schema copy) and compares field names, types, and required-ness against the deployed staging OpenAPI document — the oracle. An explicit, checked-intests/tool-endpoint-map.tsrecords the tool → operation correspondence and field renames (e.g.force → force_unsafe); a tool missing from the map, an unknown extra tool, an unclassified input field, a mapped operation absent from the engine, or a mapped body field the engine no longer has all fail loud. Runs in the PR gate (network-tolerant) and nightly (network-required). - Staging integration lane (
tests/integration/, nightly). Runs the built server as a subprocess with a freshly minted staging key and drives it over the real MCP protocol:tools/list(== 27), a read-only core loop, andaethis_decideagainst a public showcase ruleset; a negative path proves an invalid key returns a structured error result while the server stays alive. Keys are minted via the self-serve path (server-default scopes), namede2e-dx-mcp-*, and revoked + swept in teardown. staging-integration.yml— nightly + manual, report-only, emits a QA-run-shaped run record artifact for downstream ingestion; missing secrets or unreachable staging fail red, never skip-green.
- feat: authorization errors now render the server’s
hintand the missing scope readably. A403 denied_missing_permission(and401) previously printed the raw error object on commands that render their own errors (projects,whoami, …); the CLI now renders one clean line naming the missing permission plus, on its own dim line, the server’s follow-up hint (e.g. how to request access). The top-level handler and the per-command renderer now share one formatter (aethis_cli.output.format_error_detail/render_api_error), so every command surfaces the same readable message. The hint is rendered with markup disabled (so a hint containing[brackets]isn’t dropped) and non-stringmissing_permissionsitems are coerced (so a server quirk can’t turn the error into a traceback). - test: new staging integration lane (
tests/integration/, markerstaging). Acquires an API key the self-serve way (a fenced e2e user’s session → mint with the server’s default scopes, noscopesfield), drives the CLI core loop against deployed staging (whoami/status,projects list/archive,rulesets/explain/fields/decideagainst a public showcase ruleset), and asserts the negative paths a caller actually sees — a scope-reduced key’s 403 and a revoked key’s 401 — with the error envelopes checked against the machine-readable public-API contract. Report-only nightly workflow (staging-integration.yml); never gates a merge. Run locally with the one-liner intests/integration/README.md. - test: the spacecraft authoring e2e moved to its own weekly lane (
authoring-e2e-weekly.yml). It drives the LLM authoring pipeline, so it is kept out of the nightly LLM-free cadence; the model is passed explicitly viaX-Anthropic-Key, generation is bounded by an explicit iteration cap (SPACECRAFT_GENERATION_TIMEOUT), and themanualmarker stays as the local escape hatch.
Added
- Read-only cross-tenant
/api/v1/admin/*router (epic aethis-workspace#480, P1):/admin/keys,/admin/usage(daily rate-limit counters),/admin/generation-jobs(+/{id}with trace),/admin/decisions(list excludesfield_values;/{decision_id}full record),/admin/publish-audits,/admin/rulebooks,/admin/rulesets(metadata; DSL source additionally requiresrulesets:source). Gated by the new internal-onlyadmin:readscope ANDapi_key.internal == trueAND a hard-fail underDISABLE_AUTH=true(503) — scope alone is deliberately not the boundary.admin:readis never self-servable (ALLOWED_SCOPES) and never an alias target. Uniform{items, next_cursor, limit, skipped_invalid}envelope, cursor pagination, server-side max page size, default 30-day lookback on time-series lists, per-doc validation (one malformed record never 500s a list), typed filters (tenant_id=Noneonly via explicitanonymous_only=true, never a wildcard). Newadminrate-limit category enumerated in every tier (internal ≈ unlimited). Routes stay in the live/openapi.jsonbut out of the published public contract.
Added
- Decision log (epic aethis-workspace#480, P0): every
/decidecall can now be persisted server-side as aDecisionRecord(collectiondecisions), fulfilling theDecideResponsedocstring’s deferred “server-side audit persistence”. Gated byDECISION_LOG_ENABLED(default off, fail-closed) with TTL retention viaDECISION_LOG_TTL_DAYS(default 90) — the TTL index is declared in code and asserted at boot (logging disables loudly if the live index lacksexpireAfterSeconds). The write is fire-and-forget off the hot path: bounded pending-task set, client-side insert timeout, fully exception-isolated (no decision-log failure can alter a/decideresponse). Write failures log structured ERROR, increment a counter, and can alert viaDECISION_LOG_ALERT_WEBHOOK(Google Chat, rate-limited; unset = off). DecideRequest.caller_ref— optional opaque caller metadata (flat string→string dict, ≤2 KB, no$/dotted keys) stored on the decision record for the caller’s own attribution (e.g. tda-server sends{firm, application_id}). Never an authorization key or cross-principal predicate (defect shape DS-25); invalid values are dropped with a warning, never rejected.
Added
FieldDefinition(and the/schema+/decidenext_question/optimal_pathenvelope) now carries an optionalx_ui_widget: Optional[str]authoring override. Currently only"free_text"is recognised: it tells downstream consumers (Lisa’sexpected_inputemission in tda-server) to suppress the schema-derived structured-answer affordance (chips / select / date-picker) for that field and render a plain text composer, even though the field has a typedsort. Defaults toNone— purely additive. (#255, epic aethis-workspace#422)
Added
/decide: thenext_question(and eachoptimal_pathentry) now carries optionalsortandenum_valuesfields, exposing the field’s answer type (Int/Bool/String/Enum/Date/Duration) and, forEnumsorts, its allowed values. Lets callers render typed input affordances (yes/no chips, a date picker, an option list) without a second/schemaround-trip. Both default toNone, so the change is additive — existing consumers are unaffected. Mirrors thesort/enum_valuespair already on/schema’sFieldInfo. Populated on the ruleset path; rulebook/decidecallers continue to source the type from/schema. (#254, epic aethis-workspace#422)
- docs: correct stale paper citation in the construction-insurance demo. The
demo cited the withdrawn v3.6/v3.7 claim that GPT-5.4 at
reasoning_effort=lowscores 7/11 on the exception-chain subset; the paper withdrew that result in v3.8 (instrumented replication: 11/11). The demo now attributes 7/11 to GPT-5.3 only and pins the paper citation at v3.11. No code changes.
- fix(errors): attach the API’s
detail(and fullbody) toAethisAPIError. On the primary error path,classify_responseparsed the 4xxdetailonly to log it, then raisedAethisAPIError("Aethis API returned 422")— blinding callers to why the request failed. The exception message now reads"Aethis API returned 422: <detail>"when a detail is present, andAethisErrorgained.detail/.bodyattributes carrying the parsed payload (bothNonefor timeouts / connection errors). Constructor signatures stay backwards-compatible (new args default toNone). - fix(models):
DecideResponse.explanationis a single object, not a list. The field was typedlist[dict] | Nonebut the engine returnsOptional[Dict[str, Any]]({decision, groups: [...], unused_facts: [...]}), a latentValidationErrorfor any caller that actually requested one. Retyped todict[str, Any] | None. - feat(decide):
include_explanationparameter ondecide()/decide_rulebook()(sync + async). The engine has always acceptedinclude_explanationonPOST /decide, but the SDK never sent it, leavingDecideResponse.explanationpermanentlyNone. Passed through in the request payload alongsideinclude_trace; defaults toFalse. - feat(models): typed
FieldNoteandNextQuestion.notes. The engine attaches structured author guidance (note_text,source,metadata) to eachnext_question; the SDK silently dropped it. Adds theFieldNotemodel (exported from the package) andnotes: list[FieldNote]onNextQuestion, defaulting to[]so older responses without notes keep parsing. - feat(client):
list_rulesets(limit=20, offset=0)(sync + async) — wrapsGET /api/v1/public/rulesets, returning the previously-exported-but-unreachableRulesetSummarymodel. Anonymous callers get public rulesets; an API key additionally surfaces that key’s own rulesets.limitis clamped by the engine to 1-50. - docs(readme,
_base): capability-table + docstring fixes. README’s “What’s included” table now listsexplain_failure,decide_rulebook,list_rulesets,include_explanation, andFieldNote; thebuild_headersdocstring no longer names a nonexistent/next_questionendpoint.
Cross-surface review batch (aethis-mcp#50).
- Surface
next_question.notes(additive).aethis_next_questionnow renders a Notes block after the question when the ruleset author attached notes to it (each note carriesnote_text,source, andmetadata). Notes are labelled bymetadata.type(e.g.why,legal_background) when present, and each note’s text is wrapped withfenceUntrusted(...)since it is author-provided server content. Output is unchanged when no notes are present. The tool description now mentions the Notes block. - Fence
aethis_list_guidanceoutput. Theguidance_textandsourcefields returned by the server were interpolated into the tool result unfenced, unlike every sibling handler. They are now wrapped withfenceUntrusted(...)under theUNTRUSTED_PREFACEwarning (GHSA-ph7q-r9q4-922g hardening). - Send a single provider header. The per-request LLM key was sent under both
X-Anthropic-KeyandX-OpenAI-Key. It is now sent only asX-Anthropic-Key, matching howresolveLlmKeyresolves the key. - Correct the stale latency figure. Two guidance strings claimed decisions are
<5ms; corrected to<1msto match the README and the canonical figure. - Docs: rewrote the
CLAUDE.mdarchitecture section to the real layout (src/index.ts+src/client.ts+src/credentials.ts, tests undertests/) instead of the non-existentsrc/server.ts+src/tools/tree.
- fix: network errors now render one actionable line, not a raw traceback. When the API is unreachable, times out, or a DNS/TLS error occurs, every command now prints
Could not reach the Aethis API at <url>: <reason>.plus a “check your connection” hint and exits non-zero, instead of dumping anhttpxstack trace. The top-level handler catcheshttpx.HTTPError(the umbrella over connect/timeoutRequestErrors), matching the graceful handlinglogin/accountalready had. - feat: non-interactive environments bypass confirmation prompts. A truthy
AETHIS_NONINTERACTIVEorCIenv var (values1/true/yes, case-insensitive) now flips the whole process non-interactive, so destructive commands (account revoke,rulesets archive,projects archive,rulebooks archive,rulebooks tests delete) proceed without waiting on stdin, so a background job or CI step no longer hangs on a[y/N]prompt. The bypass prints a one-line notice so it’s never silently active. The explicit per-command--yes/-yflags keep working unchanged. New sharedaethis_cli.prompts.confirm_or_aborthelper. - docs: refreshed the worked examples in
decide/explain/fieldshelp to use public showcase rulesets (aethis/spacecraft-crew-certification,aethis/consumer-credit-prequalification) instead of product-specific slugs. - chore:
make installusesuv pip install -e ".[dev]"(matching the README) instead of barepip. - minor: dropped the unused upgrade-command strings from
update_check._detect_install_method(it now returns just the detected method; the concrete upgrade argv is still built byupdate’s_upgrade_argv);decidereads the decision field with a safe default so a payload withoutdecisionrenders asunknownrather than raisingKeyError.
- feat(rulebooks): declare
robot_hints:in a rulebook file and push them to the engine. Rulebook authors can now provide natural-language guidance for the conversational assistant alongside the rulebook’s other configuration.aethis rulebooks create <name> --file rulebook.yaml— a new--file/-foption reads arobot_hints:block (a sibling ofname/domain/outcome_logic) from arulebook.yaml/.jsonand sends it on create. CLI flags still ownname/domain/slug/description; only the hints are taken from the file. No--file(or a file without arobot_hints:key) is a clean no-op — behaviour is unchanged.aethis rulebooks set-logic <id> -f rulebook.yamlnow also accepts a wrapped form: when the top-level object carries anoutcome_logic:key, a siblingrobot_hints:block is pushed in the same update. A bare Expr AST file (the prior shape) is still accepted unchanged.robot_hintsis a mapping of beat-name to a natural-language string. Active beats:general_context,preamble,session_start,postamble,session_end,stuck. Reserved beats (accepted, not yet acted on):persona,conversational_style,section_transition. Unknown beat keys and non-string values are rejected client-side with a clear message before the round-trip.- New optional
robot_hintsparameter onAethisClient.create_rulebook()/update_rulebook(); omitted from the request body when not supplied, so calls against an older engine are unaffected. - Requires aethis-core with the rulebook
robot_hintsfield (aethis-core#220); mid-deploy to staging at time of writing. Against an engine without it, the field is ignored/rejected server-side.
- feat(fields):
aethis fieldsis now a command group for the full field-authoring loop. Bareaethis fields [-b <ruleset>]still shows a ruleset’s field schema (unchanged); three subcommands manage the localfields/fields.yaml:aethis fields discover— uploads the project’ssources/(creating the project if needed), runs server-side LLM field discovery, and merges the proposals intofields/fields.yamlso you start from a real draft instead of a blank file. Existing entries are preserved — only new keys are appended — so hand-authored labels/questions/hints are never clobbered. Prints the completeness score and any critical gaps. Needs an LLM key (ANTHROPIC_API_KEY), same asgenerate; without one it fails with a clear message naming the env var instead of a raw server header error. NewAethisClient.discover_fields().aethis fields pull— syncs the server’s authoritative produced fields (key + type + enum values) back intofields/fields.yamlso local matches reality after a generate. Local-onlylabel/hintsare preserved; fields absent from the server schema are kept and reported rather than silently dropped.aethis fields validate— checksfields/fields.yamlbefore upload: validtype(int/bool/string/enum/date/duration), no duplicate keys,enumrequiresenum_values. The same validation now also runs insideaethis generate, per contributing file (rulebook + ruleset), so duplicate keys within a file fail fast before any server state changes.discover/pullonly ever write a vocabulary that re-validates: an unknown server type or anenumwith no values falls back tostringinstead of producing a file the nextvalidate/generatewould reject. Writes also preserve any hand-authored keys the tool doesn’t model (e.g.description) rather than dropping them.
- feat(generate): the field spec/produced diff is surfaced after generation. After a successful
aethis generate, the CLI compares the pinned field vocabulary against what the engine actually produced and prints pinned-but-not-produced / produced-but-not-pinned fields (with a pointer toaethis fields pull) instead of the drift passing silently. - feat(init): rulesets can declare rulebook membership explicitly. A
rulebook:key in a ruleset’saethis.yaml(a path to the enclosing rulebook) now declares membership directly; the directory-position convention (<rulebook>/rulesets/<ruleset>/) remains the fallback. Theinitscaffold documents the key. - perf: source uploads are now idempotent.
discoverandgenerateshare one project-resolution + upload path, and a per-file mtime ledger in.aethis/state.jsonmeans adiscoverfollowed by agenerate(or repeated generates) only re-uploads sources that actually changed instead of re-pushing the wholesources/tree each time. - fix(generate): don’t lose the ruleset id on a fast success. The poll loop occasionally saw the job flip to
successa beat beforelatest_ruleset_idwas populated, writing a null id to state and leavingfields pull/ the field diff with nothing to work from. It now re-polls briefly for the id and only records a real one — never clobbering a prior good id with null. - example + e2e:
examples/community-grants-rulebook/is a generic rulebook (one shared field) with two member rulesets, andtests/e2e/test_rulebook_hierarchy_e2e.py(gated by themanualmarker) drives discover/validate/generate/pull against a live API and asserts the shared rulebook field propagates into both members. - No engine change required — all endpoints (
/fields/discover,/rulesets/{id}/schema,/fields/spec) are already served by aethis-core and used by the MCP server.
- feat(init): field definitions get a real home (
fields/fields.yaml).aethis initnow scaffolds afields/directory with afields.yamlfor declaring the field vocabulary (key +type+ optionallabel/question/hints). Previously fields had no dedicated home and only surfaced implicitly as theinputs:keys insidetests/scenarios.yaml.aethis generatereadsfields/fields.yaml, pins the field keys/types via the project field-spec endpoint, and routes each field’s label/question/hints through guidance so a field is defined once. - feat(init):
--kind rulebookscaffolds a rulebook.aethis init <name> --kind rulebooklays down a rulebook directory with sharedguidance/andfields/plus arulesets/directory for member rulesets. When a ruleset lives under a rulebook (<rulebook>/rulesets/<ruleset>/),aethis generatepropagates the rulebook’s guidance hints and field vocabulary into the ruleset — the rulebook definition wins on shared field keys — so a common field (e.g. date of birth) is defined once at the rulebook level and the end user is asked for it only once.--kinddefaults toruleset, so existing behaviour is unchanged.- New
AethisClient.set_field_spec()(project field-spec endpoint, already served by aethis-core / used by the MCP server). No engine change required.
- New
- feat(rulebooks list): anonymous fallthrough to the public rulebook catalogue. With no cached API key,
aethis rulebooks listnow lists the cross-tenant public catalogue (rulebooks with public visibility, active status) instead of printing the v0.19.1 pointer message — completing the parity withaethis rulesets list. A dim one-liner (“No API key — showing public rulebooks…”) distinguishes the anonymous view; with a key, the tenant listing is unchanged.- New
AethisClient.list_public_rulebooks(); use withmake_anonymous_clientso a cached key doesn’t promote the call to an authenticated tenant listing. - Requires aethis-core v0.29.0+ on the target API (live on api.aethis.ai). Against an older engine the anonymous path surfaces the server’s 401 cleanly.
- New
- fix(rulebooks list): stop prompting browser sign-in for anonymous users.
aethis rulebooks listwith no cached API key used to trigger the lazy-auth browser login — bad first-contact DX for a read-only browse command. Rulebooks are tenant-scoped, so an anonymous caller has nothing to list; the command now prints a pointer to the anonymous public catalogue (aethis rulesets list) and toaethis login, and exits 1 without ever opening a browser.- True anonymous fallthrough (listing public rulebooks without an account, mirroring
aethis rulesets list) needs engine support for a public rulebook catalogue and is tracked separately; this release removes the login prompt in the meantime.
- True anonymous fallthrough (listing public rulebooks without an account, mirroring
- feat(update):
aethis update— self-update the CLI to the latest release. Detects how the CLI was installed (uv tool, pipx, or pip) and runs the matching upgrade command.aethis update --checkreports whether a newer release exists without installing anything.- Editable (development) installs are refused with a pointer to
git pull && uv syncinstead of clobbering the checkout. - The exit-time “new release available” banner now points at
aethis updaterather than a method-specific command. - fix: the banner’s uv upgrade hint was
uv tool install --upgrade aethis-cli, which re-resolves from scratch and silently drops any extra--withrequirements (e.g. plugin packages installed alongside the CLI). Both the banner’s install-method detection andaethis updatenow useuv tool upgrade aethis-cli, which honours the original install receipt. - A successful (or no-op)
aethis updaterefreshes the banner’s 24h cache, so the notice goes quiet immediately after updating.
- Editable (development) installs are refused with a pointer to
aethis_refine now performs finding-driven incremental re-authoring: it
seeds generation from the section’s active ruleset and asks the engine for the
minimal edit to fix failing test cases while keeping passing tests green,
instead of re-authoring the whole section from scratch. aethis_generate_and_test
is unchanged (from-scratch authoring).Why this matters: fixing one wrong case in a published ruleset previously meant a
full-section regenerate — expensive, and prone to silently regressing carefully
tuned behaviour (e.g. caseworker-review criteria that intentionally yield
undetermined). Refine keeps the blast radius to the criteria that actually need
to change; the full-suite gate still guarantees no regression ships.Requires aethis-core with the mode parameter on /generate (engine ≥ the
release shipping seed-from-existing refine). Older engines ignore the body and
fall back to from-scratch generation.client.generate()/generateAndTest()accept an optionalmodeand send{mode:"refine"}on the generation request body.
- feat(refine):
aethis refine+aethis generate --mode refinefor incremental, seed-from-existing re-authoring. Instead of re-authoring a whole section from scratch, refine seeds generation from the section’s active ruleset and makes the minimal edit to fix failing tests while keeping passing tests green.aethis refine [--hint "..."] [--seed-ruleset-id <id>]— the phase-3 TDD-loop command: optionally add a guidance hint, then refine. Defaults to seeding from the section’s active ruleset.aethis generate --mode refine [--seed-ruleset-id <id>]— the same capability via a flag ongenerate;--mode fresh(default) is unchanged from-scratch authoring.AethisClient.generate()gains optionalmode/seed_ruleset_id; a no-arg call still sends no body, so it stays backwards-compatible against engines without the parameter.- Requires aethis-core with the
modeparameter on/generate(live onapi.aethis.ai). Against an older engine the flags no-op (empty body = fresh).
Add the rulebook tier to the MCP read surface. Closes
aethis-mcp#43 for the
two endpoints the engine exposes today; the public-catalogue equivalent
(
aethis_discover_rulebooks) is deferred until aethis-core ships a no-auth
rulebooks catalogue endpoint.Why this matters: until now, an agent connected via MCP could see the parts
(rulesets via aethis_discover_rulesets / aethis_list_rulesets) and
evaluate the whole (aethis_decide with rulebook_id), but had no way to
find a rulebook or inspect how its rulesets compose. Concrete failure
mode from a real 2026-05-27 session: asked whether aethis/uk-fsm was
“three rulebooks or one rulebook with three rulesets”, the MCP gave no
read path that could answer.Added
aethis_list_rulebooks— lists rulebooks in the current tenant (auth-required, tenant-scoped). Returns the fields needed to distinguish one composed rulebook from N independent rulesets:rulebook_id,slug,name,domain,status,version,outcome_logic(the composition Expr AST),ruleset_refs, timestamps. Mirrorsaethis_list_rulesets.aethis_rulebook_schema— fetches one rulebook’s composition, bridged rulesets (with names + slugs + ruleset_ids), and aggregated input fields. Accepts either a slug (aethis/uk-fsm) or an opaquerb_*id. Mirrorsaethis_schemabut at the rulebook tier.AethisClient.listRulebooks()/getRulebookSchema()— wrapGET /api/v1/public/rulebooks/andGET /api/v1/public/rulebooks/{slug-or-id}/schema. The schema helper preserves the literal/in slugs (soaethis/uk-fsmhits the engine’s{namespace}/{name}matcher) and URL-encodes opaque ids.
Deferred
aethis_discover_rulebooks— the cross-tenant public catalogue equivalent ofaethis_discover_rulesets. The engine’s/api/v1/public/rulebooks/endpoint is tenant-scoped + auth-required on prod today; no anonymous catalogue variant exists. Will land once aethis-core adds it.
- feat(output): gh-style machine-readable output mode (
--output json,--json fields,--jq). Every list/show command (and the decision commands) now emit structured JSON on demand, soaethis rulesets list --output json | jq '.[0].slug'just works instead of trying to scrape ANSI-coloured Rich tables.--output table|json— pick the format. Default:tableon a TTY,jsonwhen piped (matches gh’s pipe-friendly autodetect).--json FIELDS— implies--output json; takes a required comma-separated value (--json id,name) that limits the payload to those fields. (gh’s bare---jsonintrospection trick is not yet exposed — Click/Typer’s option parser can’t cleanly distinguish “flag with no value” from “flag followed by positional”, so it’s deferred to a future--list-fieldsflag.)--jq EXPR— pipe JSON output throughjqbefore printing. Requires thejqbinary on PATH; clear error with install hint if missing.- Commands migrated:
rulesets list/show,rulebooks list/show/get-fields/tests list/schema/explain/decide,projects list/show,account keys,profile list,guidance list,fields,explain,decide,status. Each command has a sensible JSON shape —status --output json | jq .identity.key_idreturns the live key id without rooting through any prose. - Footer hints (
Try: aethis ...) are suppressed in JSON mode so pipes get clean output. - New module
aethis_cli/render.pyis the single emit point; new test filetests/test_render.pycovers the matrix.
- breaking(guidance export):
--outputrenamed to--output-fileto avoid clashing with the new global--outputflag. Short form-ounchanged. Affects scripts that pipe to a named file:aethis guidance export --output foo.yaml→aethis guidance export --output-file foo.yaml(or-o foo.yaml).
- fix(status, whoami): read the multi-profile credentials file the same way every other command does.
aethis login --api-key ...writesprofiles.<name>.api_keyto~/.config/aethis/credentials(the multi-profile schema introduced in v0.10), butaethis statusandaethis whoamihad stale local resolvers that only looked for a flat top-levelapi_key(andwhoamiwas looking at the wrong filename,credentials.yaml). Result: after a freshaethis login,aethis statusreportedno API keyandaethis whoamireportedNo Aethis API key configured, even though the same key worked foraethis projects list,aethis generate, and every other authoring command.- Both commands now route through the canonical
resolve_cached_key()helper inauth_helpers.py, which honoursAETHIS_API_KEYenv → active profile → keychain → legacy.yamlfile. - The
_resolve_cached_keysymbol is renamed toresolve_cached_key(public). The legacy_resolve_key_silent(status_cmd) and_resolve_api_key_lax(whoami_cmd) are removed. - Regression test in
tests/test_status_cmd.pywrites a real multi-profile credentials YAML to a tempXDG_CONFIG_HOMEand asserts both commands surface the key.
- Both commands now route through the canonical
- chore(server): tighten MCP instructions against decision extrapolation. Adds a “Reporting decisions” section to the server
instructionsblock (visible to every client model as part of its system prompt on connect). New rules forbid asserting facts that are not in the tool response, generalising a single-ruleset decision to a composite outcome, naming rulesets/rulebooks not yet observed in the session, and offering follow-up calls against unverified slugs. Triggered by a real user trace where a model summarising auk-fsm-child-eligibilitydecision closed with an offer to run the broaderaethis/uk-fsmrulebook “to get the complete household-level decision” — the rulebook exists but currently 422s on prod (emptyruleset_refs, see aethis-core#90), so the offer overstated what would actually happen. Advisory, not enforced — but client models reliably honourinstructionsblocks.
- feat(explain-failure):
Aethis.explain_failure()+AsyncAethis.explain_failure()— wrapsPOST /api/v1/public/rulesets/{ruleset_id}/explain-failure, returning the failing criterion and a targeted fix hint for a mismatched/decideresult. Acceptsfield_values,expected_outcome("eligible"|"not_eligible"|"undetermined"), and an optionaltest_name(default"test"). Return type isdict[str, Any]to matchexplain()/get_source()— can be tightened once the response shape stabilises. Note:ruleset_idmust be the concrete identifier (not a slug); the underlying endpoint does not currently resolve slugs. Previously, callers had to drop to rawhttpxfor this endpoint — flagged inrecipes/evaluate-a-case.mdxandrecipes/debug-a-decide.mdx.
- docs(readme): rulebook surface advertised on the PyPI landing page. The v0.5.0 release shipped
decide_rulebook()andrulebook_idonDecideResponse, but the README still framed the SDK as ruleset-only. Adds a dedicated “Composed rulebook” section with a runnable UK FSM example, the always-scope-gated note, and the async equivalent. - docs(install): switch
pip installtouv addper workspace no-pip rule. The PyPI landing page is a public-facing surface bound by.claude/rules/no-pip.md. Addsuv pip installas a venv-friendly alternative. - docs(engine_version): update sample audit-field comment from
aethis-core@0.10.0toaethis-core@0.27.0— matches live prod engine. - docs(beta): clarify that decision endpoints are anonymous only for single rulesets — rulebook decide is always scope-gated, so the SDK’s “anonymous when no key” claim needed a footnote.
- feat(rulebook):
Aethis.decide_rulebook()+AsyncAethis.decide_rulebook()— evaluate a composed multi-ruleset rulebook through the SDK. Mirrorsdecide()but sendsrulebook_idin the payload. Accepts either an opaquerb_<id>or a slug (e.g.aethis/uk-fsm). Requires an API key — rulebook evaluation is always scope-gated. Closes #14. Requires aethis-core v0.27.0+ live on the target API for slug-form rulebook paths. - feat(models): add
rulebook_id: Optional[str]toDecideResponse— surfaces the rulebook identifier when the response was a composed-rulebook decide. Backwards-compatible: ruleset-only decides keeprulebook_id=None.
- docs(readme): v0.27.0 accuracy pass. Three fixes for fresh-developer accuracy:
- Documented
rulebook_idas an alternative toruleset_idonaethis_decide— mutually exclusive; composed-rulebook evaluation always requires an API key. - Quickstart example: corrected field name from
speciestospace.crew.species(the actual field ID in the spacecraft-crew-certification ruleset). - Windsurf config path: corrected from
.windsurf/mcp.jsonto~/.codeium/windsurf/mcp_config.json(canonical path per aethis-cli README).
- Documented
Add rulebook surface to
aethis_decide — closes the converged-2-term
client-completeness gap for MCP. The tool now accepts either
ruleset_id (single ruleset) or rulebook_id (composed rulebook),
mutually exclusive. Mirrors aethis-sdk-python v0.5.0 and
aethis-cli rulebooks decide.Added
aethis_decidetool acceptsrulebook_idas an alternative toruleset_id. Pass an opaquerb_<id>or a slug likeaethis/uk-fsm. Composed-rulebook evaluation is always scope-gated by the engine — anonymous callers get HTTP 401.AethisClient.decideRulebook(rulebookId, fieldValues, options?)— parallel todecide(); sendsrulebook_idin the/decidepayload.
Changed
aethis_decidedescription and schema updated to reflect both paths. Tool validates that exactly one ofruleset_id/rulebook_idis provided.
Requires
- aethis-core v0.27.0+ live on the target API for slug-form rulebook
paths. The
rulebook_idbody field on/decidehas been supported since aethis-core v0.18.x.
- docs(readme): v0.27.0 accuracy pass. Three fixes for fresh-developer accuracy:
- Install block: removed the
pip installfallback (uv and pipx are the recommended forms per workspace policy). Development section:pip install -e ".[dev]"→uv pip install -e ".[dev]". - Added Rulebooks command-group section documenting the converged 2-term model surface shipped in v0.14.0–v0.16.1 (
aethis rulebooks+aethis rulesetspromote-to-live). - Updated engine_version example to
aethis-core@0.27.0(was absent; clarified to current production version).
- Install block: removed the
- docs(rulebooks set-logic): the docstring example for
field_ref.keynow matches engine behaviour. Phase A.16 (aethis-core v0.26.0+) added per-section aggregate group synthesis, sofield_ref.key = <ruleset_name>resolves to the AND of that ruleset’s groups. The unscoped group-name and scoped<ruleset_name>.<group>forms remain available for advanced compositions. Requires aethis-core v0.26.0+ live on the target API.
- feat(rulebooks):
aethis rulebooks set-logic— set the composition expression on a rulebook. The composition expression (server fieldoutcome_logic) is an Expr AST that combines per-ruleset outcomes into the rulebook’s final decision. Previously settable only via raw PATCH; now exposed via the CLI for multi-ruleset rulebooks (e.g. UK FSM’schild_eligibility AND (household_criteria OR universal_infant)).aethis rulebooks set-logic <id> -f logic.yaml— load from YAML/JSON fileaethis rulebooks set-logic <id> --logic '<json>'— inline JSON- Exactly one of
--file/--logicis required; both forms reject non-object payloads at the client side so server validation isn’t the first line of defence.
- feat(rulesets): ruleset lifecycle commands scoped to a rulebook. Phase B.1b of the converged 2-term model. Adds four new sub-commands under
aethis rulesets:aethis rulesets list <rulebook>— list rulesets in a rulebook (grouped byruleset_namewith version counts, live version, and observed states). The legacy-p <project_id>and--publicmodes are preserved while the project-scoped authoring pipeline retires in a future phase.aethis rulesets create <rulebook> <ruleset_name> [-n "Display name"]— create a new draft Ruleset inside the rulebook. The display name auto-derives fromruleset_nameif not provided (child_eligibility→Child Eligibility).aethis rulesets show <rulebook> <ruleset_name>— full version history for one ruleset name (bundle_id, version, state, created), with live version highlighted.aethis rulesets promote-to-live <rulebook> <ruleset_name> <ruleset_id> [--note "..."]— atomically promote atesting-state ruleset version tolivevia the Phase A.4 service. Auto-cuts a new rulebook version; previous live ruleset is archived.
- feat(client): four new
AethisClientmethods —create_ruleset_in_rulebook,list_rulesets_in_rulebook,show_ruleset_in_rulebook,promote_ruleset_to_live. - Requires aethis-core v0.20.0+ live on the target API (Phase A.8 endpoints).
- feat(rulebooks): new
aethis rulebookscommand group. First user-facing surface for the converged 2-term authoring model (workspace PR #64, aethis-core PRs #133-139). A Rulebook is the whole form — the execution unit — that owns a locked field vocabulary, composition logic, rulebook-level test cases, and an integer version history.aethis rulebooks list— list tenant rulebooksaethis rulebooks show <id-or-slug>— full configurationaethis rulebooks create <name> --domain <d> [--slug ...]— create draftaethis rulebooks set-fields <id> -f fields.yaml— replace locked vocabularyaethis rulebooks lock-fields <id>/unlock-fields <id>/get-fields <id>aethis rulebooks tests add <id> -f scenario.yaml— embed full-form test caseaethis rulebooks tests list <id>/delete <id> <tc_id>aethis rulebooks activate <id>/archive <id>— lifecycleaethis rulebooks decide <id> -i '{...}' [--explain]— evaluate composed rulebookaethis rulebooks schema <id>/explain <id>— combined schema + explanations
- feat(client): new
AethisClientmethods for every rulebook REST endpoint (create / list / show / update / activate / archive / set-fields / lock-fields / unlock-fields / get-fields / add-test / list-tests / delete-test / decide-rulebook / get-rulebook-schema / explain-rulebook). - Requires aethis-core v0.19.0+ live on the target API (the Phase A.6 endpoints).
- The legacy
aethis projects/aethis generate/aethis test/aethis publishcommand tree is unchanged in this release — replacement lands in the next minor (Phase B.1b: ruleset lifecycle + project retirement). No backward-compat shims are planned past public release.
- feat(models): add
name: Optional[str]to ruleset response models — surfaces the human-readable section name introduced in aethis-core v0.18.0. AddsRulesetSummary(anonymous catalogue /GET /api/v1/public/rulesets) andRulesetListItem(project-scoped /GET /api/v1/public/projects/{id}/rulesets) as typed models, and adds the samenamefield toSchemaResponse. Backwards-compatible: pre-backfill rulesets serialise withname=None.
Add optional
name parameter to aethis_publish tool — lets clients
override the human-readable section name when publishing a ruleset.
Companion to aethis-core v0.18.0’s PublishRequest.name field.Added
aethis_publishtool now accepts an optionalnameparameter. When supplied, it overrides the section name stored on the ruleset (default is a titlecase ofsection_id, e.g."english_language"→"English Language"). Section names are surfaced in rulebook responses so end users can see which sections compose a rulebook.AethisClient.publish()now accepts a thirdname?: stringargument and includes it in the POST body when set.
Tests
server.test.ts— two newaethis_publishcases: forwardingnameto the client and echoing it in output; confirmingnameis omitted when not provided.client.test.ts— two newpublish()cases: body containsnamewhen provided; body is absent when neitherlabelnornameis set.
Surface the human-readable section
name in aethis_list_rulesets and
aethis_discover_rulesets tool output. The engine has been returning
name on both RulesetSummary (public catalogue) and RulesetListItem
(project-scoped) responses since aethis-core v0.18.0; the MCP server
already forwarded every API field verbatim via JSON.stringify, so the
data was reaching the LLM, but the tool descriptions didn’t advertise
the field. The descriptions now mention name so models know to read
and surface it to users (e.g. “Knowledge of language and life in the
UK” instead of just b_123…).Changed
aethis_list_rulesetstool description now mentions the human-readablenamefield returned alongside ruleset ID, status, version, field count, and rule count.aethis_discover_rulesetstool description now listsnamein the documented response shape.
Tests
aethis_list_rulesetsandaethis_discover_rulesetsserver tests assert thenamefield passes through to the LLM-facing JSON output.
- feat(rulesets): show the human-readable section
namecolumn inaethis rulesets listoutput (both the public showcase and project-scoped tables). Surfaces the new field from aethis-core v0.18.0.
- feat: pluggable auth providers. Profiles now carry an optional
auth_mode(default"api_key") andaudiencefield. The newaethis_cli.auth_providersmodule exposes a process-local registry; plugins (e.g.aethis-cli-internal) canregister_provider("gcloud_id_token", ...)to add staff/internal auth schemes without touching the published package.AethisClientaccepts an optionalauth_providercallable, andmake_authed_client(...)picks the right provider based on the active profile’s mode. - feat:
aethis statusnow prints the active profile name + auth mode (plus audience when set). For non-api_keymodes it shows “provider-minted at request time” instead of calling/me, which is X-API-Key-only. - chore: un-hide the
--base-urlglobal flag inaethis --help(it was already implemented, justhidden=True).
Security hardening pass. Bundles the v0.5 security review fixes into one
release. Closes #33, #34, #35; addresses GHSA-ph7q-r9q4-922g (disclosed
on publish).
Security
- GHSA-ph7q-r9q4-922g (high) — prompt injection via unsanitised API
response text in
aethis_explain_failure.formatExplainFailurenow wraps every API-supplied free-text field (diagnosis,dsl_hint, criteriontitle/rule_text/source_refs) in an<api_response>fence and prepends a one-line preface telling the model the contents are data, not instructions. Literal closing tags inside payloads are neutralised so a payload cannot break out of the fence. The samefenceUntrustedhelper has been applied to other free-text API surfaces (aethis_next_question,aethis_discover_sections,aethis_refine_sections,aethis_discover_fields,formatTestResults). - #33 —
src/credentials.tsnow resolves the credentials file viafs.realpathand asserts the canonical path sits under$HOME(or under an absoluteXDG_CONFIG_HOMEthe user controls); refuses withUnsafeCredentialsErrorotherwise. Permissions check now matchesssh/aws-clibehaviour: any group/other bit set on the credentials file → refuse withPermissions 0NNN ... too open. Run: chmod 600 <path>. - #34 —
progress_detailfrom the polling API is sanitised before it lands on stderr: control characters are stripped (TAB preserved) and the body is capped at 120 visible chars +…. Full-fidelity output is gated behindAETHIS_MCP_VERBOSE=1. Prevents the server from injecting terminal escape sequences or PII into anything that captures the MCP process stderr.
Changed
- #35 — Authoring tools now accept safer per-call key forms.
- New
anthropic_key_env: string— name of an env var the MCP server reads at call time. Preferred. The raw value never appears in the tool call so it does not land in the MCP host’s session transcript JSONL. - New
anthropic_key_keychain: string— macOS keychain reference, eitherservice:accountor justaccount(service defaults toaethis-anthropic-key). - Raw
anthropic_key/openai_keyarguments remain accepted for backwards compatibility but are now marked[sensitive — do not echo or log]in the schema; deprecated in tool descriptions. resolveLlmKey(exported fromsrc/credentials.ts) consolidates the resolution chain and throwsMissingLlmKeyErrorif every form is empty.- Applies to
aethis_generate_and_test,aethis_refine,aethis_discover_fields,aethis_refine_fields,aethis_discover_sections,aethis_refine_sections.
- New
Docs
- README: new “Passing your Anthropic key safely” section showing env / keychain forms first; raw-key form marked deprecated.
- CLAUDE.md: new gotchas covering safe-key resolution and the untrusted-content fencing helper.
- fix(decide):
aethis decide --explainno longer crashes withAttributeError: 'str' object has no attribute 'get'. The CLI previously treated the engine’sexplanationfield as a flatlist[dict], but the public decide route returns a layered{decision, decision_path?, groups: [{group, status, criteria: [{title, status, supporting_facts?, ...}]}], unused_facts}shape. The “Rules” block now walks the actual structure and renders each group + criterion with PASS/FAIL marks, supporting fact field/value pairs underneath satisfied criteria, and a final list of unused fields (provided answers that no satisfied criterion referenced — useful for catching field-name typos).
- fix(login): default
AETHIS_CLERK_CLIENT_IDto the OAuth Application registered on theclerk.aethis.aiClerk instance. The previous default belonged to a different Clerk app, soaethis loginreturnedinvalid_clientagainst the dev-tools domain set in 0.12.1. - fix(account): default
AETHIS_CLERK_DOMAINtoclerk.aethis.aiforaethis account generate(matching the 0.12.1 change toaethis login); previously still pointed at the immigration domain.
- feat:
decide,explain, andfieldsno longer prompt for sign-in when no API key is present. Public rulesets are now accessible with zero setup — the CLI silently uses an anonymous client and lets the server return an error only if a private ruleset is requested. - fix: hide
--base-urlglobal flag fromaethis --help(internal dev override;AETHIS_BASE_URLenv var unchanged) - docs: reorder
aethis --helpto lead with the no-auth explore flow, then authoring
- docs: surface
aethis-skillsas the optional agent workflow layer on top of MCP.
- fix: default Clerk domain changed from
clerk.aethis.legaltoclerk.aethis.aiso developer portal users can authenticate viaaethis login(closes aethis-cli#40)
- fix: align
package.jsonrepository metadata with GitHub provenance so npm Trusted Publishing can verify the package source.
- fix: pin
zodto v3 so the MCP SDK tool registration types match the build-time schema shape;npm publishnow runs theprepublishOnlyTypeScript build successfully.
- fix: update
examples/session.pyto useAETHIS_RULESET_IDenv var (was deprecatedAETHIS_BUNDLE_ID) and replace stale internal default with the publicaethis/construction-all-risksslug
- docs: fix stale
bundle/bundle_id/aethis_create_bundleterminology indocs/demo-construction-insurance.mdanddocs/agentic-decision-systems.md— these files were not caught by the v0.3.0 rename sweep. All references now useruleset/ruleset_id/aethis_create_ruleset - chore: bump
server.jsonversion to 0.4.1 (was lagging behindpackage.json) - security: regenerate
package-lock.json— bumpshono4.12.12 → 4.12.18,fast-uri3.1.0 → 3.1.2,ip-address10.1.0 → 10.2.0,postcss8.5.8 → 8.5.14; clears all 6 open Dependabot alerts (closes #19)
- feat: new
aethis_discover_rulesetstool — lists the cross-tenant public showcase catalogue (no authentication required). Mirrors the no-auth policy ofaethis_decide/aethis_schema/aethis_explain. Returns slug, ruleset_id, description, field_count, rule_count for each entry; the slug or ruleset_id can then be passed to the existing decision tools. Distinct fromaethis_list_rulesets, which remains tenant-scoped and authenticated. Tool count: 24 → 25. - feat:
client.discoverRulesets(limit, offset)wrappingGET /api/v1/public/rulesets. - docs:
aethis-decideprompt and server-instructions now point ataethis_discover_rulesetsfor first-time discovery (no key) before falling back toaethis_list_projects→aethis_list_rulesetsfor authenticated tenant browsing.
- fix: remove
examples/demo_core.sh(internal dev script referencingaethis-coreby name and a private API path — not intended for public release) - fix: update
tests/e2e/test_spacecraft_e2e.pyto resolve the spacecraft fixture fromexamples/spacecraft-crew-rules/instead of an internal path; drop internal service name from comment - docs: fix “rule bundle” → “ruleset” in
examples/spacecraft-crew-rules/README.md
- feat(updater): gh-style update-check banner. On startup the CLI
kicks off a background thread that queries PyPI; if a newer release
is available it prints a one-line notice to stderr at exit:
“A new release of aethis-cli is available: 0.11.0 → 0.12.0 — to
upgrade, run: <method-aware command>”. Detects whether the install
came via uv tool, pipx, or pip and renders the matching upgrade
command. Result is cached for 24 h at
~/.config/aethis/update_check.json. Suppressed automatically when stderr is not a TTY (CI, piped output). Disable withAETHIS_DISABLE_UPDATE_CHECK=1. The check never blocks the command — failures are silent.
- feat(rulesets):
aethis rulesets list --publiclists the cross-tenant public showcase catalogue (no auth required). When run with no--project-idand no project context, falls through to the public catalogue automatically with a one-line hint — so a fresh signup sees something the moment they install the CLI instead of an empty list. Combine withaethis fields -b <slug>/aethis explain -b <slug>/aethis decide -b <slug>to fully exercise a ruleset without an API key. - feat(profiles): named credential profiles with both per-invocation
flag (
aethis --profile new-dev …) and sticky default (aethis profile use new-dev). Manage withaethis profile list/use/add/remove. Reserved profile nameanonymousforces unsigned mode — handy for testing what a fresh signup sees without losing your admin key.aethis login --profile <name>writes into the named slot. Credentials file format upgraded to{active_profile, profiles: {...}}; legacy single-key files are read transparently and rewritten to the new shape on next save. - feat(client):
AethisClient(unsigned=True)andmake_anonymous_client()helper for paths that must hit the anonymous surface without accidentally sending a cached key. - feat(client):
client.list_public_rulesets(limit, offset)wrappingGET /api/v1/public/rulesets.
- feat(publish): thread
--forcethrough to the server-side TDD gate introduced inaethis-core0.11.0.client.publish()gains aforce_unsafe: bool = Falsekeyword;aethis publish --forcenow passesforce_unsafe: truein the request body so the server-side gate is bypassed (and apublish_force_bypassaudit event is recorded). Older engines ignore the field — no breakage. Without--force, the new gate refuses publishing over a failing test suite even when the CLI’s own test gate is bypassed (e.g. by a direct curl that doesn’t use the CLI). Closes the cli/server asymmetry that nearly shipped a 10/11 ruleset to a canonicalaethis/*slug on 2026-05-07.
- docs: link to the test-driven authoring guide on docs.aethis.ai and surface the publish-gate guarantee (rulesets cannot publish with a failing test) in the private-beta callout. Reference surface only — no code changes
- docs: surface the test-gate guarantee —
aethis_publishrefuses to publish a ruleset with a failing test, derived from positioning bible §5/§7. Strengthens the existing Note to an Important callout and annotates the publish line in the four-stage workflow - docs: drop
force=truemention from troubleshooting — surfacing the override on the public README undermines the “cannot be published with failing tests” guarantee. The API parameter remains in the engine; whether to deprecate it is tracked separately - docs: fix tool count (25 → 24); tools table sums to 24 (5 + 7 + 8 + 2 + 2). Fixed in README header and in CLAUDE.md
- docs: link to docs.aethis.ai/agents/onboarding from Install section
- docs: link to docs.aethis.ai/agents/onboarding from MCP one-liner section
- docs: remove positioning paragraph above Install — reference surface (per aethis.os/positioning/surface-types.md); the tagline is enough
- docs: add private-beta callout for authoring endpoints (decision endpoints remain anonymous)
Changed
- docs: align README with positioning bible — add problem/solution/methodology intro paragraph before Install section.
- docs: add
aethis-bible:markers to derived copy blocks (sourced frompublic-messaging.md §3/§4). - fix: terminology audit found no deprecated “rule bundle” or
<5msinstances in README; no replacements needed.
Changed
Aethis(api_key=...)andAsyncAethis(api_key=...)now acceptapi_key=None(or no argument) for the developer beta. Evaluation endpoints (/decide,/schema,/explain,/source) work anonymously, so the SDK no longer forces a key on instantiation. Whenapi_keyis omitted, thex-api-keyheader is simply not sent. Authoring endpoints will still return 401 without a key. Existing callers passingapi_key="..."are unaffected.- README quickstart now shows
Aethis()(no key) as the primary form, targetsaethis/uk-fsm/child-eligibility(a live public ruleset) instead of the datedeng_lang:20250912-ec5d7c23, and prints the audit fields (inputs_hash,decision_id,decision_time,engine_version) added in 0.3.2. Configuration table updated:api_keyis now documented as optional during the developer beta. examples/oneshot.pyrefreshed to match: no key required by default,AETHIS_BUNDLE_IDenv var renamed toAETHIS_RULESET_ID(catching the 0.3.0bundle → rulesetrename it had missed), targets the live UK Free School Meals ruleset, prints the audit fields.
Notes
- Backwards-compatible:
Aethis(api_key="ak_live_...")continues to work exactly as before. - This pairs with the public-surface positioning that evaluation is free during the developer beta — see
docs.aethis.ai.
Added
DecideResponse.decision_id— per-call audit identifier returned by the engine.DecideResponse.inputs_hash— canonical SHA-256 fingerprint of the input set.DecideResponse.decision_time— ISO-8601 timestamp of the decision.DecideResponse.engine_version—aethis-core@<semver>string identifying the engine that produced the decision.
Fixed
DecideResponsepreviously declaredruleset_idtwice; Pydantic silently overrode the first declaration with the second. Deduplicated.- The four audit fields above were already returned by
/api/v1/public/decidebut were silently dropped by Pydantic because the model didn’t declare them. Callers can now read them directly off the typed response — no need to reach for the raw JSON. This is the audit-trail fingerprint that the docs and homepage prominently advertise (inputs_hash,decision_id); shipping an SDK that hid it was a defect.
Notes
- Backwards-compatible. All four new fields default to
None, so older engines that don’t emit them still parse cleanly.
Fixed
aethis_sdk.__version__now resolves from installed package metadata viaimportlib.metadatainstead of a hardcoded constant. Previously reported"0.1.0"on every install regardless of the actual package version. Falls back to"0.0.0+unknown"only if the package is imported without being installed (editable dev or zip-on-PYTHONPATH).- Package description on PyPI:
"…and bundle schemas"→"…and ruleset schemas"to match the v0.3.0 public-surface rename.
Added
- README PyPI / Python-version / License shields.
- docs: restructure README as dev MCP docs — Install / Quick start / Tools / Setup leads, positioning sections (Problem, Accuracy, When to use this, How it works, Example walkthrough) removed; their content belongs in docs.aethis.ai or the benchmarks repo
- docs: trim narrative paragraphs across Quick start, Conversational eligibility, and Authoring; collapse repeated Tips into terse callouts
- docs: header tagline rewritten to a single factual line; link bar updated to the new structure
- docs: normalise tone to documentation register — replace argumentative Proof section with one-liner accuracy claim, neutralise example framing, trim sales-y bullets in When to use this
- docs: add private-beta callout for authoring tools (decision tools remain public, no key required)
- docs: align README with positioning bible — promote 225-scenario accuracy framing
- docs: add aethis-bible: markers to derived copy blocks
- docs: fix latency claim to <1ms (was <5ms)
- fix: replace deprecated “rule bundle” terminology with “ruleset”
- docs: remove Why Aethis section — package README is a reference surface (per aethis.os/positioning/surface-types.md); install / quick start / authentication is the right lead, not a problem statement
- docs: add private-beta callout for authoring tools (decision tools remain public, no key required)
- docs: clarify in Authentication that aethis login requires an invite during the beta
- docs: align README with positioning bible — add Why Aethis section, solution framing, TDD methodology beat
- docs: add aethis-bible: markers to derived copy blocks
- fix: replace deprecated “rule bundle” terminology with “ruleset” in pyproject.toml description
Changed (Breaking)
- Renamed the public bundle concept to ruleset throughout the SDK to match the
aethis-core 0.10.0API contract. Everybundle_idparameter and JSON key is nowruleset_id. URL paths inside the client moved from/api/v1/public/bundles/...to/api/v1/public/rulesets/.... TheSessionconstructor now takesruleset_idand exposessession.ruleset_idinstead ofsession.bundle_id. Class names:BundleSummary→RulesetSummary.
Required
- Engine
aethis-core 0.10.0or newer. Older engines respond at the legacy/bundles/*paths and this client will 404. Pinaethis-sdk==0.2.0to keep working against an older engine.
- Breaking: renamed the public bundle concept to ruleset throughout the MCP tool set, to match the
aethis-core 0.10.0API contract. The compiled rule artefact is now called a ruleset in every tool name, parameter, and prose description. Specifically:- Tools:
aethis_create_bundle→aethis_create_ruleset,aethis_list_bundles→aethis_list_rulesets,aethis_archive_bundle→aethis_archive_ruleset - Parameters: every
bundle_id→ruleset_id - JSON keys returned to the agent:
bundle_id/latest_bundle_id/bundle_version/deprecated_bundles/result_bundle_id/bundle_refs→ruleset_idetc. - URL paths inside the client:
/bundles/...→/rulesets/...
- Tools:
- This release requires
aethis-core 0.10.0or newer. Older engines respond at the legacy/bundles/*paths withbundle_idJSON keys; this client expects/rulesets/*and will 404. Pinaethis-mcp@0.2.6if you need to keep working against an older engine until you can deploy. - MCP tool renames are part of the public LLM-facing contract. Coding agents that have learnt the old tool names (
aethis_list_bundlesetc.) from training data will get “no such tool” errors and need to retry against the new names. Tool descriptions explicitly call out the new naming so the LLM picks it up on first read.
- Breaking: renamed the public bundle concept to ruleset throughout the CLI to match the
aethis-core 0.10.0API contract. The compiled rule artefact is now called a ruleset everywhere — in command names, in flag names, in JSON keys, and in prose. Specifically:aethis bundles list/archive→aethis rulesets list/archive--bundle-idflag →--ruleset-idclient.list_bundles()/archive_bundle()/get_bundle_schema()/explain_bundle()/get_bundle_source()/set_bundle_visibility()SDK methods →*_ruleset- JSON keys
bundle_id/latest_bundle_id/bundle_version/bundle_refs→ruleset_idetc. - Default scope strings
bundles:read/explain/write→rulesets:*(validated against the engine’s permission registry)
- This release requires aethis-core 0.10.0 or newer. Older engines return
bundles:*scopes and the CLI will reject them as invalid. Pin toaethis-cli==0.7.2if you need to keep working against an older engine until you can deploy.
- Docs: replaced two stale
aethis.ai/sign-uprequest-access pointers in the README authoring section withaethis.ai/developer-access. After the Clerk cutover,/sign-upserves the Clerk SignUp form for invitees rather than the Notion request-access form. No code or behaviour changes.
- Docs: replaced the stale
aethis.ai/sign-uprequest-access link withaethis.ai/developer-accessin the README “Author your own rules” section and in theaethis whoamihint shown when the active key has no authoring scope. After the Clerk cutover,/sign-upserves the Clerk SignUp form for invitees rather than the Notion request-access form, so external “Request access” pointers were broken. No code path changes.
- Docs: README Quick start now leads with
aethis mcp install --target all(via aethis-cli v0.5.0+). The manualclaude mcp addand per-client JSON tabs are demoted to “Manual install” beneath. Setup section gains a Keys & security subsection coveringAETHIS_API_KEYvsANTHROPIC_API_KEYplacement (MCP client config, not shell), rotation workflow (aethis account generate+aethis account revoke), and multi-machine guidance. - Discoverability:
package.jsonkeywordsextended withregulation,policy,eligibility-check,deterministic-decision— matches the highest-intent search terms used by developers in regulated domains. Existing keywords retained. - CLAUDE.md updated to note the
aethis mcp installinstall path so future contributors don’t re-document the manual JSON as primary.
- Docs: README gains a dedicated Authentication section explaining the three modes (
aethis loginfor explicit setup, lazy auth for inline mid-command sign-in,--no-promptfor CI). Authoring quickstart leads withaethis init(the v0.7.0 wizard prompts for a name and runs sign-in itself, soaethis loginas a separate step is no longer needed). Environment-variable table expanded to coverAETHIS_BASE_URLandANTHROPIC_API_KEY. Troubleshooting entry forAuth errornow mentions the lazy-auth prompt and--no-prompt. CLAUDE.md updated to document theaethis mcp installpath, lazy-auth helper, and--no-promptflag for future agents working on the CLI. No behaviour change.
- New:
aethis initfirst-run wizard. With no args, prompts for the project name (default = current directory name); a positionalaethis init <name>keeps working unchanged. If no API key is cached, triggers the same OAuth flow asaethis loginbefore any filesystem writes — Ctrl-C during browser sign-in no longer leaves a half-scaffolded project on disk. After scaffolding, prints the next-step ladder (aethis sections discover→fields discover→generate --poll) so new users have a clear path forward. New--no-promptflag for scripted use; with that flag, missing required values fail fast and missing auth surfaces a cleanAuthRequirederror instead of opening a browser. 10 new tests covering prompted, non-prompted, no-auth + interactive, no-auth +--no-prompt, and name-validation paths. Closes #15.
- New: lazy auth. Authenticated commands (
aethis projects list,generate,publish, etc.) now detect missing credentials or 401 responses and offer an inline browser sign-in prompt:"No API key. Open browser to sign in? [Y/n]". On accept, the same OAuth flow asaethis loginruns, the key is cached, and the original command retries — exactly once, no infinite loops. Non-TTY stdin/stdout (CI, pipes) and the new--no-promptglobal flag skip the prompt and surface a cleanAuthRequirederror.--api-key <key>still bypasses the helper entirely. New helper moduleaethis_cli/auth_helpers.py; the OAuth flow insidecommands/login_cmd.pywas factored into a reusablerun_browser_login(). 17 new tests intests/test_lazy_auth.py. Closes #12.
- New:
aethis mcp install --target <client>writes the MCP server entry into your editor’s config in one shot. Supportsclaude-code(project-level.mcp.json),cursor(~/.cursor/mcp.json),claude-desktop(~/Library/Application Support/Claude/claude_desktop_config.jsonon macOS,~/.config/Claude/...on Linux),windsurf(~/.codeium/windsurf/mcp_config.json), and--target allfor everything at once. Idempotent, preserves any other configured MCP servers.aethis mcp uninstall --target <client>reverses the install. Closes #16.
- UX:
aethis login --helpnow reads “Sign in and store an API key locally. First-time setup — this is all you need.”aethis account generate --helpclarifies it’s for additional keys (rotation, multi-machine, scoped access). After successfulaethis login, a tip line points ataethis status/aethis account keys. README quickstart collapses any “first login then generate” sequence into a singleaethis loginstep. No behaviour change. Closes #13.
- Docs: README install section now leads with
uv tool install aethis-cli(recommended) andpipx install aethis-cli, withpip installin a venv as the third option. Pairs with Aethis-ai/docs#12. Closes #14.
First version published to npm since 0.2.2. The
v0.2.3 tag exists in git but predates the publish workflow — it never reached npm. This release rulesets all work since 0.2.2.Registry
- MCP Registry submission ready. Added
mcpName: io.github.aethis-ai/aethis-mcptopackage.jsonand a top-levelserver.jsondeclaring the npm package, transport, and environment variables. Submit viamcp-publisherafternpm publish.
Breaking Changes
openai_keyparameter renamed toanthropic_keyonaethis_generate,aethis_generate_and_test, andaethis_refine. The old parameter name is still accepted for backwards compatibility but will be removed in a future release.
Improvements
- Better error messages on generation failure. Failed jobs now surface classified error details (invalid key, rate limit, connection failure) instead of “unknown error”.
- Sends both
X-Anthropic-KeyandX-OpenAI-Keyheaders for backwards compatibility with older API versions. aethis_explain_failureclarification. Tool docs now note thatruleset_idmust be the concrete ID from a/decideenvelope; slugs are not yet resolved on this endpoint (tracked in aethis-core#51).
Docs
- Proof section updated to cite the Simpson et al. 2026 benchmark paper. Replaced the pre-paper 11-scenario table (GPT-5.4-mini 82%, GPT-5.3 27%) with paper-backed figures from Table 8b of the published benchmark. Removed the 27% GPT-5.3 claim — the paper identifies that figure as a harness-configuration bug; the corrected value is 63.6%.
- Proof section: add §6.10 LegalBench external-validation paragraph. v3.8 of the paper adds external validation across 9 LegalBench tasks (949 held-out cases). Combined paired-binomial McNemar’s: p < 0.001 vs Sonnet 4.6, p = 0.003 vs Opus 4.7, p < 0.001 vs GPT-5.4. Linked to the public LegalBench harness at
confidently-wrong-benchmark/legalbench/. - Proof section: replaced 11-scenario subset table with v3.8 adversarial extension (§6.4.1). The v3.7 11-scenario exception-chain table no longer differentiates current frontier models from the engine (GPT-5.4 default and low both 11/11, Opus 4.7 11/11). The Proof section now leads with the v3.8 adversarial extension (20 newly-authored scenarios; engine 20/20; Opus 4.7 18/20; GPT-5.4 default 19/20 with 0 reasoning tokens; Sonnet 4.6 19/20) and the shifting-ground argument from paper §6.5 Finding 6.
- Use
aethis/construction-all-risksslug in CAR proof example for stable URL across ruleset regenerations. - Invite-only beta messaging replaces “rolling out now” framing throughout README — explicit approval-gated framing aligned with current onboarding.
docs.aethis.aibadge added to README.
Internal
- Added
.github/workflows/publish.yml(provenance via OIDC +NPM_TOKEN) so future tag pushes auto-publish. - Added Claude PR review workflow (dry-run mode).
- Added internal
CLAUDE.mdfor agent onboarding.
Two bug fixes that block the documented quickstart against public bundles.
Bug fixes
aethis decide -b <slug>/explain -b <slug>/bundles archive -b <slug>now accept slugs. The classifier in_id_utils.classify_idpreviously returned"unknown"for slugs (e.g.aethis/uk-fsm/universal-infant), andrequire_bundle_idrejected them with"is not a valid Bundle ID". The public API resolves both bundle IDs and slugs on/decide,/schema, and/explain, so the CLI now passes both through. Error message updated to mention slugs and link toaethis bundles list.aethis fields -b <bundle>no longer requires anaethis.yaml. It now uses the sameload_client_or_fallback()helper asdecide,explain,bundles, andprojects— read-only commands work from any directory. Previously this command errored out with"No aethis.yaml found"even when called with a concrete bundle reference.
Added
DecideResponse.slug— stable, human-readable handle for the ruleset (e.g.aethis/uk-fsm/child-eligibility). Set when the resolved ruleset was published under a slug;Noneotherwise. Prefer this overruleset_idfor any reference that should survive ruleset regeneration.SchemaResponse.slug— same handle, surfaced fromGET /rulesets/{id}/schema.
Notes
- Backwards-compatible. Existing code that reads
ruleset_idkeeps working unchanged;slugis purely additive. - Requires the
aethis-coreengine release that surfaces the field in/decideand/rulesets/{id}/schemaresponses (rolling out 2026-04). Older engines will simply leaveslug=None.
aethis status output polish
- Server line now shows just the URL when it’s the default (
https://api.aethis.ai) — the(default — no override)suffix was noise in the common case. Overrides (AETHIS_BASE_URL,aethis.yaml) still show source with a green marker. - Identity line now says
✗ API key rejected (run \aethis login` to re-authenticate)when/mereturns 401/403/404, instead of the raw✗ 404 from /me (Not Found)` HTTP message. Other HTTP errors keep a contextual message.
This release ships the rich-status and read-only-from-anywhere work that the 0.2.0 notes already described but which hadn’t actually been merged into a published release yet. (The code was sitting in a local branch; the prior 0.2.x/0.3.x wheels still had the minimal status command.)
aethis status — context-aware summary
aethis statuswith no args now prints CLI version, resolved server URL (with source — env / yaml / default), loadedaethis.yaml, bundle id from.aethis/state.json, andwhoamiidentity (key id, tenant, tier, scopes,can_author). Helps answer “what will my next command actually hit?” before running it.aethis status -p <project_id>(or from inside a project dir) still shows generation progress, appended after the global summary.
Read-only commands usable from anywhere
aethis explain,decide,bundles list,bundles archive,projects list,projects show,projects archiveno longer require anaethis.yamlin the current directory — they fall back toAETHIS_BASE_URL(or the defaulthttps://api.aethis.ai).aethis explain/decidenow reject Project IDs (proj_*) passed to-b/--bundle-idwith a one-line hint pointing at the Bundle column ofaethis projects list, instead of silently 404’ing.
Internals
- New
resolve_base_url_with_source()/load_client_or_fallback()helpers inaethis_cli/config.pythat the above commands share. - New
aethis_cli/commands/_id_utils.py+ test coverage for bundle-id validation. - New tests for
explain,status, and_id_utils.
Docs cleanup
- README and docs.aethis.ai/interfaces/cli no longer document
AETHIS_BASE_URLor showbase_url:in theaethis.yamlexample — public users always hithttps://api.aethis.ai, and the documented values were just duplicating the default. The env var still works as an override for devs and CI; it’s intentionally undocumented. - Dropped the
AETHIS_CLERK_DOMAINenv var from the README (marked “development only” and confusing for public users). The override still works in code.
Trim public CLI to the developer API surface
The public CLI now only ships commands every developer can use againsthttps://api.aethis.ai. Privileged and staff-only commands have been removed and will live in a separate internal plugin package.Breaking changes:- Removed
aethis source— internal-only DSL viewer; moved to theaethis-cli-internalplugin. - Removed
aethis account permissions— IAM permission registry; internal-only. - Removed the
aethis guidance domain …group (and the deprecatedaethis domain guidance …alias) — domain-level guidance is staff-managed. - Removed the global
--base-urlflag (plus the per-command--base-urlonlogin,account generate,account keys,account revoke). TheAETHIS_BASE_URLenv var still overrides the default. The flag had no meaning for the public API target and cluttered--help.
- The CLI now discovers plugins via Python entry points under the
aethis_cli.pluginsgroup. A plugin exposes one callableregister(app: typer.Typer) -> Noneand attaches extra commands to the root app. Plugin load failures print a single warning to stderr and never crash the CLI. - The staff-facing
aethis-cli-internalpackage uses this hook to re-attachsource,domain guidance,permissions, and the--base-urlflag.
Consolidated guidance command tree
aethis domain guidance ...moved underaethis guidance domain ...— thedomaingroup exists only to hostguidance, so having two top-level trees for the same concept was confusing. All four subcommands (add,list,import,export) behave identically on the new path.- The old
aethis domain guidance ...path still works as a hidden deprecated alias: invocations continue to succeed and emit a one-line deprecation notice to stderr. It is no longer shown inaethis --help. Planned removal in a future release.
aethis status — global CLI context
- New behaviour:
aethis status(no args) now prints a one-screen summary of the current CLI context: CLI version, resolved server URL (with source —--base-url/ env / yaml / default), loadedaethis.yaml+ project, bundle id from.aethis/state.json, and whoami identity (key id, tenant, tier, scopes,can_author). Answers “what will the next command hit?” — the usual cause of “why is my project missing?” is talking to the wrong server. - Backward compatible:
aethis status -p <project_id>(or invoked from a project dir) still shows generation progress, now appended after the global summary.
UX improvements for read-only commands
aethis explain,decide,bundles list,bundles archive,projects list,projects show, andprojects archiveno longer require anaethis.yamlin the current directory — they fall back toAETHIS_BASE_URL(or the defaulthttps://api.aethis.ai) when invoked from anywhere.aethis explainanddecidenow reject Project IDs (proj_*) passed to-b/--bundle-idwith a one-line hint pointing at theBundlecolumn ofaethis projects list, instead of silently proceeding to a 404.aethis --base-url <url>is now a top-level flag, equivalent to settingAETHIS_BASE_URLfor one invocation. Lets you hit staging or a self-hosted instance without editingaethis.yaml.aethis projects listprints a short tip after the table showing how to copy a Bundle value intoaethis explain -b ….- Configuration and authentication errors now render as a single red line via the existing
cli()handler, not a Rich traceback panel.pretty_exceptions_enable=Falseis set on every Typer app.
Better --help
- Top-level
aethis --helpnow shows common flows (status, list, explain, decide), authoring flow, and how to target a different server. explain,decide,bundles list,projects list, andstatusall have “Examples:” blocks in their per-command help.
New Tools
aethis_add_domain_guidance— Add cross-section guidance hints at domain level (e.g.uk_citizenship). Applies automatically to all projects in the domain during generation.aethis_list_domain_guidance— List all active domain-level guidance hints.aethis_list_guidance— List all guidance hints accumulated for a project. Use before adding new guidance to avoid duplicates.aethis_explain_failure— Diagnose a failing test case. Returns criterion statuses with DSL metadata and a targeted fix hint.
Improvements
aethis_add_guidancenow acceptsprocess_type("rule_generation"|"field_extraction"). Usefield_extractionfor field design principles (solicitor navigation, raw-facts principle). Defaults to"rule_generation".aethis_add_domain_guidanceacceptsnotes— SME commentary or legislation provenance stored on the hint. Never sent to the LLM.- Two-level hint retrieval: generation now fetches domain-level hints (cross-section) alongside project-level hints in a single pass.
New Tools
aethis_discover_fields— Discover input fields from source text. Returns field names, types, and completeness assessment. Call before writing test cases.aethis_refine_fields— Iterate on field discovery with targeted feedback.
Improvements
- Added
aethis-authorandaethis-decideMCP prompts for compatible clients (Claude Desktop, Cursor, VS Code Copilot).
Initial release.
Features
- Decision tools:
aethis_schema,aethis_decide,aethis_next_question,aethis_explain - Discovery tools:
aethis_list_projects - Authoring tools (TDD workflow):
aethis_create_ruleset,aethis_generate_and_test,aethis_add_guidance,aethis_refine,aethis_publish,aethis_archive_project,aethis_archive_ruleset - HTTPS enforcement for remote hosts
- Exponential backoff with retry on 429/502/503/504
- Works with Claude Desktop, Claude Code, Cursor, and Windsurf
Note: v0.1.0 usedaethis_create_ruleset(renamed toaethis_create_ruleset) andaethis_project_status(replaced byaethis_list_projects). These tools were removed in v0.2.x.
Initial release.
Features
- Account management:
aethis account generate(browser OAuth),aethis account keys,aethis account revoke - Project authoring:
aethis init,aethis generate --poll,aethis test,aethis publish - Decision tools:
aethis decide,aethis fields,aethis explain - Project management:
aethis projects list,aethis bundles list,aethis bundles archive - Security: HTTPS enforcement, OS keychain storage, PKCE OAuth flow
- Example: Spacecraft Crew Certification Act 2049 with 5 golden test cases