Changelog
All notable changes to this project are documented in this file. The format is based on Keep a Changelog, and the package adheres to Semantic Versioning. The data contracts (core-world, exposure, extraction, connection, risk, agentic, ambiguity, search, temporal, and broker schemas) are versioned independently of the package; see DATA_DICTIONARY.md.
Unreleased
Added
- Agent-spy PII-to-identity evidence. The experiments catalogue now records the corrected, locally frozen external-consumer study of detector-error propagation into identity resolution, relationship inference, and exposure assessment over SynthWorld 0.17.0. The evidence page preserves the run identity, independent metric denominators, isolation and determinism checks, limitations, and local-only artifact status without presenting the result as a SynthWorld benchmark, product claim, or independently reproducible release.
0.17.0 - 2026-08-21
Added
- Generated enterprise-agentic standard and longitudinal tiers (#27). A new independently versioned V2 configuration and artifact family adds deterministic multi-organisation standard scale plus a 180-day longitudinal schedule covering repeated credential rotation, joiner/mover/leaver changes, suspension, policy activation, agent offboarding with active credentials, revocation propagation, evidence retention, and audit-time scoring. The released smoke V1 and frozen Asteria/C08 contracts remain unchanged. Public topology, opaque credential-handle, and lifecycle projections stay physically separated from evaluator-only case labels and truth; generic loaders, CLI config/public-only export, derived denominated integrity metrics, end-to-end observed-action scoring, and an external digest-bound runtime/memory receipt complete the generated-scale path. Stress remains explicitly deferred to generic scale issue #3 and generated tiers remain workloads rather than vendor leaderboard claims or mandatory package fixtures.
- Enterprise agentic identity policy pilot. A repository example now compares experiment-owned RBAC, ABAC, ReBAC, and default-deny combined policy views over one generated enterprise-agentic smoke world. The policy runner consumes only the verified public tree and writes digest-bound decision traces; a separate scorer verifies the matching evaluator tree and emits independent metrics plus deterministic, visibly watermarked comparison HTML. A new end-to-end guide covers generation, public projection and layout JSON, the evaluator overlay, Explorer HTML, replay, scoring, reproducibility, and the boundary between this teaching pilot and SynthWorld authorization contracts or production enforcement.
- Generated enterprise-agentic Explorer adapter (#149). The
synthworld visualizecommand now renders a verified released generated enterprise-agentic smoke package as deterministic, self-contained HTML when selected explicitly with--package-profile generated-enterprise-agentic. An independently versionedenterprise-agentic-generated-v1projection records the generated configuration digest, world identity, tier, and public artifact-set digest; a new1.0.0generated layout contract computes a deterministic kind-layered grid from the projection alone, so no frozen1.0.0Explorer or Asteria contract widens in place and published Asteria HTML bytes are unchanged. The public path consumes only the verified public tree; evaluator truth requires the separate evaluator tree, is digest-bound through the reused overlay contract, and renders visibly watermarked. Unsupported generated tiers or package versions fail explicitly, and the shared #52 renderer, assets, and interaction are reused without a second viewer implementation. - Packaged Asteria Explorer renderer (#52). The
synthworld visualizecommand now renders the checksum-verified published Asteria Agentic v1 public package as deterministic, self-contained HTML with an interactive authority graph, chain inspection, and timeline replay. Cytoscape and build-time ELK versions are pinned, generated assets and layout coordinates are checksum-bound, no runtime network resources are loaded, and dependency notices ship in the wheel. The new layout2.0.0contract explicitly records world seed, world schema version, and visualisation profile identity while preserving layout1.0.0unchanged. Evaluator truth requires a separately verified evaluator package and produces prominently watermarked output; public HTML contains no evaluator overlay. Generated enterprise-agentic support is supplied by the separately versioned adapter described above; enterprise authorization packages remain outside Explorer.
0.16.0 - 2026-08-17
Added
- Stable enterprise authorization consumer API (#139). A curated
synthworld.enterprise.consumernamespace now exposes the authoring rows, vocabularies, limits, artifact/result wrappers, compilers, prediction rows, evaluators, and exact canonical digest helpers needed for a released-wheel workflow. A versioned operator-side compiler provenance artifact maps canonical topology and directory-policy locations to compiled opaque IDs without entering public product input or evaluator truth. Authorization public bundles now carry their evaluation scope, and an isolated-wheel test exercises topology import, all three mechanism families, composition, public-only prediction, and scoring. - Publicly solvable composed enterprise authorization scoring (#137). New
independently versioned public evaluation-scope, system submission, and
evaluation-report contracts score effective and final decisions separately
from RBAC, ABAC, ReBAC, conflict, binding, lifecycle, and runtime-gate
outcomes. Exact public cell inventories, artifact and schema bindings,
deterministic system/policy metadata, and per-metric denominators are
enforced without changing existing
1.0.0schemas or golden bytes. - Discriminating adversarial enterprise authorization cases (#138). A new independently versioned reference profile adds hidden single-factor pairs for tenant, scope, credential binding, time, clearance, and RBAC/ReBAC authority composition. Its public policy supports generic tenant inequality without enumerating negative target IDs, candidate action attempts remain separate from persistent grants, and opaque attempt IDs reveal no pair or verdict labels. Independent metrics report total cohorts and discriminating denominators; seven weak baselines fail their dedicated dimensions without changing frozen schemas or golden artifacts.
- Published Phase 1 and Phase 2 enterprise authorization evidence. The GitHub release carries checksum-bound historical, reproduction, and reference archives from the Topaz experiment conducted against SynthWorld 0.15.0. The signed release source records the expected digests and the documentation states which archives contain evaluator evidence. These retained artifacts do not retroactively rescore the experiment with the new 0.16.0 contracts.
0.15.0 - 2026-08-15
Added
- Generated enterprise-agentic smoke vertical (#27). A new independently
versioned configuration and Python API deterministically generates a bounded
fictional organisation with humans, logical agents, runtimes, opaque
credentials, resources, ownership, attenuated delegation, revocation,
attribution, and authorised/adversarial action cases. Construction routes
through the hardened base agentic projection, emits checksum-bound separate
public/evaluator trees, derives denominated topology and integrity metrics, and
can be selected explicitly with
generate-enterprise-agentic --profile generated. Public-only and complete artifact-root loaders now reject inventory, canonicalization, checksum, cross-binding, semantic, derived-metric, and declared generator drift; dedicated CLI tasks validate and score external generated traces. An isolated-wheel consumer test and event-replay guidance preserve the oracle boundary. Existing fixed enterprise smoke and frozen Asteria/C08 bytes are unchanged; standard and longitudinal generated tiers remain follow-up work. - Explorer v0.1 public projection contracts. A preview Python API projects the published Asteria Agentic v1 public package into deterministic graph and timeline records, with independently versioned evaluator overlays and layout manifests. Digest binding, UTC ordering, acyclic compound-node validation, collision-safe UUID5 identities, typed collection properties, and explicit public/evaluator separation are enforced. Interactive Cytoscape/ELK rendering, CLI integration, additional benchmark adapters, and generated-scale support remain deferred.
0.14.0 - 2026-08-11
Added
-
Independent C08 evidence-binding v2 contracts. Asteria and enterprise lineages now have separately versioned public input, evaluator truth, submission, report, manifest, deterministic reference-generation, and evidence-quality metric contracts. Public inputs expose opaque binding handles and same-kind distractors while exact required bindings remain in evaluator truth. Existing v1 contracts and frozen bytes remain unchanged.
-
C08 v2 candidate registry metadata. Repository-local metadata records the Asteria and enterprise C08 v2 benchmark identities as candidates with pending publication gates. These entries do not publish either benchmark externally or claim external hosting, viewer support, download availability, or re-download verification.
-
Independent C08 v2 frozen-artifact candidates. Asteria now commits an exact five-file root/public/evaluator manifest tree; enterprise commits an exact four-file root-manifest/
SHA256SUMStree with a packaged fail-closed loader and fixed-seed identity comparison. Both public contracts use opaque binding handles plus same-kind distractors instead of exposing a unique kind-to-ID answer. Independent manifest schemas, expanded v1 hash locks, explicit offline report scope, and exactly two metric-only discrimination records accompany the candidate bytes. Native adversarial findings and repository verification gates are recorded as resolved. Nothing is externally published or deployed. -
A curated benchmark registry now records independent lifecycle, benchmark-kind, evaluation-mode, artifact-sensitivity, integrity, and publication-gate evidence for every current benchmark family.
-
A source-complete Blume documentation site. The repository now carries a structured public documentation tree, generated capability and benchmark catalogues, dark-preview builds, documentation-impact routing, and distribution audits. The site is ready for a separately authorized deployment but is not published by this release.
-
Guarded Hugging Face publication planning. A local publication manifest, offline validator, and protected dry-run workflow bind publication authority to the resolved benchmark registry. Uploads remain disabled: no benchmark or operation is authorized and the outstanding external gates remain explicit.
-
A safely fictional EADS-shaped adapter example. The repository-only, humans-only example accepts bounded JSON or YAML fixtures, emits separately typed public and evaluator projections, and records deterministic path-bound provenance. It is not shipped in the wheel, performs no network access, and makes no real-EADS compatibility claim.
-
Future enterprise authority designs. Reviewed design records reserve non-conflicting C15/C16 v1 and Face B contract, schema, package, benchmark, and manifest identities. They define canonical normalization, collision rejection, UUID5 domain separation, public/evaluator boundaries, and staged gates without implementing or freezing those deferred families.
Changed
- Agentic evidence metrics and deployment-pattern coverage now state their reporting-only limits, metric polarity, denominators, null and abstention semantics, and declaration-versus-observation boundary without changing metric values, scoring versions, schemas, or frozen artifacts.
- Benchmark, documentation, release, and Hugging Face governance now fail closed on same-version frozen-inventory additions, compare push transitions with the pre-push commit, audit the exact documentation distribution, and publish the already-verified release artifact rather than rebuilding it.
0.13.0 - 2026-08-06
Added
- A deterministic enterprise identity and access surface, with authorization
evaluated against standards-shaped projections (#7, #27). SynthWorld now
generates a fixed enterprise identity/access universe — organisations, units,
populations, principals, accounts, groups, roles, permissions, entitlements,
and opaque authorization targets — and evaluates access against it through
three bounded oracles: a directory RBAC oracle with role and group closure,
a bounded ABAC oracle over declared attributes, and a bounded ReBAC
oracle over relationship tuples. Access derivation is an implementation detail
of the oracle rather than a published topology, so no artifact describes the
model as an identity topology.
New modules:
enterprise,enterprise/rbac,enterprise/abac,enterprise/rebac,enterprise/authorization,enterprise/conformance,enterprise/identity_fabric,agentic/enterprise. - Vendor-neutral projections so a real authorization engine can consume the
world without a bespoke adapter.
enterprise/projections/emits an RFC 7643-style SCIM user/group projection, an OpenFGA authorization model and relationship tuples for the bounded ReBAC subset, and an AuthZEN 1.0 request projection with per-field provenance. Each projection carries an explicit mapping profile recording what it can and cannot represent, so a projection gap is declared rather than silently lossy. A Shared Signals/CAEP mapping profile is declared, but temporal event emission is deliberately deferred and no SET envelope is constructed yet. - Two smoke benchmarks and a graded assurance ladder. An enterprise identity
fabric smoke benchmark and an enterprise agentic smoke benchmark publish public
inputs and evaluator truth under
enterprise-identity-access-contract/; a contextual access benchmark profile publishes undercontextual-access-contract/.continuous_assuranceaddssmoke,standard,longitudinal, andheld_outtiers governing assurance cadence. Note these are assurance tiers, not generated-world scale tiers — theenterprise_agenticscale ladder #27 asks for remains open atsmokeonly. - An executable agent-authority run protocol.
agent_authorityandassuranceadd a staged run protocol with signed execution receipts, component provenance, and evidence claims, so an evaluation run is reconstructible from its receipt rather than trusted on assertion. - The ambiguity v2 pack gets its difficulty from a computed error floor, not a
codebook (#80). The v1-style surfaces encoded each identity index in cleartext, so a
~30-line normaliser recovered every relation and scored 1.0000; two successor designs
fell to pool inversion. The fix follows the reviewed plan: each kind draws a base from
a pool of confusable clusters (
Sorensen/Sorenson/Soerensen, a transposed phone, a swapped day/month),EQUAL/NEARshare the base whileFARredraws from a stationary mixture, and every value passes through one structured-noise operator applied identically under every relation. Identity recovery stays free and expected; the relation is carried by overlapping distance distributions. The pack publishes its genie floor — the Bayes error of the generator, estimated with a stated N and Wilson interval and keyed to a digest of every decision-relevant constant — plus the enumerated channel invariants (kernel stationarity, an identical one-value marginal under every relation, a per-base sibling-landing mass gate, form bijectivity and constant cross-form distance, and an artifact-factorization check) and a gated technique premium, so real resolution technique is rewarded rather than anti-taught. New modules:ambiguity_evidence,ambiguity_surfaces,ambiguity_channel,ambiguity_floor, with v2 serialization/metrics/baselines support andexamples/compute_ambiguity_floor.py. - Agent-authority and contextual-access receipt builders now seal honest
failed-run receipts: a failed product execution produces a manifest with
execution_status=failedandevaluation_status=not_evaluatedthat binds only the product-stage artifacts and never loads evaluator truth, so an assurance corpus can no longer be structurally biased toward successful runs. Receipt validation enforces the paired statuses and the product-only artifact inventory, and run manifests whose evidence claim is not supported by the systems under test (live-lab claims with reference-only components) are rejected. Managed-service provenance innot_exposedobservability states additionally forbids evidence references, and the contextual execution receipt leavesstimulus_digestunset because that lineage executes the public input directly. - Contextual-access receipts now expose the same explicit two-phase live-run boundary as agent-authority receipts. External runners may finish and attribute the product stage before constructing completion metadata; the finalizer then replays the plan, public input, adapter, component inventory, provenance, and artifact digests before evaluator truth is loaded. Existing deterministic contextual receipt bytes remain unchanged.
- An opt-in disposable agent-authority reference deployment now executes the public enterprise-agentic smoke world across isolated Docker networks. It produces live observation-v2 receipts covering L01-L06, the exact declared L07 baseline/SUT inventory, and measured/unsupported L08 targets, while keeping runtime credentials in a destroyed named volume and scanning canary and token markers out of receipts, logs, and container metadata. A new two-phase receipt finalizer lets live runners record completion metadata only after external execution without exposing evaluator truth before product output is durably staged.
- Agent-authority observation schema
2.0.0corrects L06 clock semantics without changing the frozen observation-v1 schema. It records one explicit monotonic revocation epoch, non-negative acknowledgement offsets, and signed send/completion offsets so pre-revocation in-flight requests are representable. Receipt validation dispatches v1/v2 observations and binds them to scoring formulas1.0.0/2.0.0; migration guidance forbids guessing offsets from ambiguous v1 rows. - The 12-case authority-change governance conformance fixture from #73 is now an
additive frozen benchmark. Its public and evaluator payloads remain physically
separate; their visibility manifests and exact raw bytes are path-bound by a
packaged
SHA256SUMS, verified by the packaged loader API, regeneration tests, and isolated-wheel checks. No existing golden bytes changed. - The broker-removal pack is scored through the unified evaluator:
evaluate_broker_removalprojects each family’s headline ratio into the standardEvaluationReport, the CLI gainssynthworld evaluate broker, andexamples/evaluate_broker_adapter.pyis the worked Idcognito-style adapter #5 asked for - public timeline in, versioned assessment out, scored against regenerated truth. Closes the last acceptance criteria of #5. - Propagation lag is representable and scored (#65). Downstream copies carry their own
removal tick (
Nonenever goes), a newslow_propagationcase has copies that catch up late rather than never, a submission can predict the completion tick, andpropagation_lagreports mean absolute error with support. The credulous baseline now predicts completion at confirmation - “done means done everywhere” - and eats a 14-tick error on exactly the case built to price that claim; the example adapter’s modest grace period cuts it to 4.
Changed
- Ambiguity grammar
2.0.0, v2 schema2.1.0.render_relation/render_valuedelegate to the structured-noise channel; the old_SPACE/_surfacecodebook is gone.Relation.EQUALno longer means “byte-identical” — it is one value transcribed once per record, rendered identically only with probabilitysigma— and the charter docstrings that claimed otherwise are rewritten. The v1 pack and its frozen artifacts are untouched and stay byte-identical. display_namerendersfamily, given(#86). The two name kinds are scored as separate evidence, so the boundary between them must be readable off the value; the oldgiven familylost it whenever a pool entry carried a space. No rendered name contains", ", so the split is unambiguous.- Temporal schema
1.2.0.ListingTruth.downstream_refs(bare strings) becomesdownstream_copieswith per-copy removal ticks, and a recorded reappearance must now coincide with a publishedLISTING_REAPPEAREDevent - truth that disagrees with the public timeline is refused as corrupt input. Deliberately asymmetric with removal, which stays unpinned because a confirmation is the broker’s claim and the phantom case exists to show the claim can be false.BrokerAssessmentmoves to1.1.0for the new prediction field; propagation state is now read as of the assessed tick, so a slowly propagating deletion no longer scores identically to one that never propagates. DenominatedMetricmoved tosynthworld.models, below the evaluation/partition import cycle it was about to create;ambiguity_partitionre-exports it unchanged.
0.12.0 - 2026-08-04
Changed
-
Breaking (ambiguity answer key):
same_name_and_date_of_birthis nowinsufficient, notseparate(#77). The pair is two people in canonical truth, but the public evidence — matching name and birth date, nothing distinguishing them — cannot justify concluding it. The pack’s first consumer abstained on exactly this pair and independently gave the same reason. Membership truth is unchanged; the public and memberships artifacts are byte-identical, and only the dispositions artifact was re-cut (new digest inGOLDEN_REVIEW.md). A resolver scored against the old key that answeredseparatehere was being rewarded for clairvoyance; one that abstains is now scored correctly. -
EVALUATION_SCHEMA_VERSIONis0.2.0.TaskMetricgains optionalfamilyandsupport_meaning, so every task’s report carries two more keys - extraction, entity resolution, relationship inference and risk included, even though none of their metrics changed meaning. The wire shape is what moved, so the schema knob is what moves; no per-task scoring version changes, because a scoring version here means the metric definitions and those are untouched. A stored0.1.0report does not load under the new model — the report’sschema_versionis a single-value literal — so read archived reports with the library version that wrote them. -
Agentic metrics are grouped into five families and every denominator says what it counts, so a report can be read by family and each ratio re-derived rather than trusted. No metric value moves. The split carrying the most information is
observabilityagainst the rest: an agent can decide well and record badly, or the reverse, and those need different fixes. Measured on the reference trace, wrecking the recording drops observability to 0.25 while identity resolution, authorization and delegation stay at 1.0; wrecking the decisions drops authorization to 0.40 while observability stays at 1.0. -
AGENTIC_BENCHMARK.mdgains a per-metric glossary: what each measures, what 0.0 and 1.0 mean, and its denominator. Sixteen of the twenty metrics had no mention in any top-level document, so the only way to learn what they measured was to read the scorer or diff scores between policies. (They were cited inagent-authority-contract/control-catalogue.yamland its design-intent notes, which a first version of this entry overlooked while claiming a repository-wide count.)
0.11.0 - 2026-08-03
Security
- Nine of eleven known channels through which the ambiguity pack’s answer key was recoverable from its public artifact are closed. Two remain open and are described below; regenerating a pack with this version does not make it safe against those two. Anyone who generated evaluation packs with an earlier version should still regenerate: nine channels are a great deal worse than two, and a system under test could otherwise reach the right answer without doing the task at all, so scores measured against those packs do not mean what they appear to. The frozen canonical pack was affected too and is re-cut here, with new digests recorded in GOLDEN_REVIEW.md.
- Eight of the eleven were metadata bound to the label — collection ordering, name pools indexed by a scenario ordinal, positional record identifiers, source types constant per scenario, repetition counts, attribute counts, a distinctive locality token, and cross-listing multiplicity. Each is closed by deriving the value from the evidence rather than from the case.
- The ninth was different and is the reason for the new
keyparameter: the substitution plan was a deterministic function of a published seed over canonical values that live in public source, so it could be recomputed and inverted rather than correlated. Rebuilding it recovered the disposition on 0.929 of pairs against a 0.467 baseline.generate_ambiguity_variantnow requires akeythat is never serialized; passUNKEYEDto reproduce the published packs, or at least 16 bytes fromsecrets.token_bytesfor evaluation. - The two that remain open, tracked in
#68. Non-ASCII display names
appear in only one scenario, so a search for them identifies a
mergepair on every seed measured. Source-type agreement impliesseparateon every pair measured where the two sources match. A key closes neither, because neither depends on a draw: they are properties of a fixed case list, which #62 addresses. Treat scores on the affected scenarios accordingly until then.
Changed
BROKER_SCORING_VERSIONis2.0.0. The scoring formulas changed rather than grew: four families moved from an assessed-listings denominator to the discovered world, a removal request counts as warranted only when the system itself concluded the listing is the subject’s, andrequest_correctnessbecamerequest_recallbecause its denominator was always recall’s. The same submission scores differently, so two reports at one version would be incomparable.TEMPORAL_SCHEMA_VERSIONis1.1.0. Additive: the public artifact gainedlistingsand truth gainedattributable, both defaulted, so every1.0.0artifact still parses and a consumer that ignores the new field reads what it did.
Added
-
A deterministic temporal slice for privacy exposure, and a broker deletion-and-reappearance pack scored on top of it. Virtual time is an integer tick, never a wall clock;
materialisereturns the events at or before a tick, so a system is asked what it knew when it could have known it. Seven named cases - six failure modes and a clean-removal control: a clean removal, a phantom removal the broker confirms but never performs, a reappearance after genuine removal, reseller copies surviving a source deletion, a refusal, a listing that was never the subject’s, and a stale binding after a move. The clean and phantom cases emit the same sequence of event kinds at the same ticks, so the hardest one cannot be read off the timeline. Replay refuses histories that cannot happen — a confirmation with no request, a reappearance with no removal — while admitting repeated requests and conflicting statuses, which are cases rather than corruptions. -
Public listing content and a published subject identity, so attribution is answerable rather than guessable. A first revision emitted lifecycle events with no content at all and never said who the subject was, which left the listing that is not theirs indistinguishable from the six that are. Content is drawn from one vocabulary and every readable page carries the same attribute kinds, so neither the attribute count nor a distinctive token substitutes for reading the values.
ListingTruth.attributablemarks the case whose page carries a common name and nothing to corroborate it: declining is correct there and deciding is unwarranted, the same distinctionPairDisposition.INSUFFICIENTandSearchMatchTruth.INSUFFICIENT_EVIDENCEdraw. -
evaluate_broker_assessment, reporting six families that are never combined: discovery, identity matching, request correctness, completion, propagation and recurrence. Every score publishes its numerator, denominator and the denominator’s meaning. Two reference policies run in CI, gated on properties rather than numbers: neither may resolve the pack, both must overstate propagation, and recurrence must separate them — trusting broker confirmations catches no reappearance, while continuing to watch catches every one and still cannot see the phantom removal or the surviving copies. -
Scoring for the oracle-free search projection:
evaluate_search_judgementsseparates false accepts from false rejects and from unwarranted decisions on results the public text cannot settle, reports coverage beside precision so abstaining everywhere cannot look perfect, and reports distinct findings against accepted results so a consumer that fails to collapse syndicated copies is visible. Errors are broken out by difficulty tier with that tier’s support, because raw error counts rank tiers by size rather than by failure rate.SearchMetricsandSearchEvaluationpublish their denominators, ascoring_versionand a task discriminator, matching the ambiguity evaluation channels, andSearchTruthBundlenow records the seed it describes. -
Separate ambiguity membership and evidence-disposition evaluation channels. A complete
EntityResolutionPredictionis validated and scored directly against explicit membership truth with denominated pairwise and B-cubed metrics; a public-only projection derives forced binary decisions for the selected task pairs without discarding the raw partition or loading truth. -
A consumer-integration publication boundary:
.local-assurance/is excluded from Git and package builds, repository tests reject private consumer symbols, dependencies, unreviewed adapter paths, and force-added local artifacts, and a named CI check plus code-owner review protects the boundary-defining files. Contribution guidance keeps one-off product execution out of public CI. -
A
synthworld validate agentic-tracecommand that checks an observed-action JSONL submission for structural and cardinality correctness before scoring, with no access to evaluator truth. It reports every bad row in one pass with line numbers, exits0for valid and1for invalid, and prints a human summary by default or a machine report with--json. The guarantee is one-directional: a valid result meansevaluate agenticwill not raiseEvaluationInputError. It is deliberately stricter in one case, rejecting a submission in which every row is empty, because the scorer would accept that and award a perfectleast_privilege_accuracy. -
load_public_agentic_bundle, which loads and checksum-verifies a public-only Asteria tree without reading any evaluator artifact, plusTraceValidationReport,TraceValidationIssue, andvalidate_trace_jsonl. -
An asserted
patternon the timestamp property of the published trace schemas.formatis an annotation rather than an assertion in JSON Schema 2020-12, so the schemas previously accepted a naive timestamp, a non-UTC offset, and"not-a-date"— all rejected by the model. -
A
schemastarget inmake cirunninggenerate_trace_schema.py --check, so schema drift fails the build, and an isolated-wheel check exercising the new command’s accept and reject paths. -
A design-intent trace per agent-credential pattern class in
agent-authority-contract/examples/, generated bytools/generate_design_intent_traces.py, with assumptions and a scored coverage table indocs/design-intent-assumptions.md. These are explicitly not measurements: no implementation was run and no product was tested. What they show is each pattern’s observability ceiling — most usefully that short-lived scoped credentials match proxy injection on decisions and temporal correctness while scoring zero on delegation provenance, attribution and accountable ownership, because those are directory facts rather than token claims. -
An adapter template in
agent-authority-contract/adapter-template/that reads only the public package, runs as shipped to produce a structurally valid trace, and isolates the integration work in one function. -
tests/test_trace_schema_agreement.py, asserting that the models and the published schemas accept the same bytes across a mutation corpus, with the two known pydantic coercion divergences declared explicitly. Addsjsonschema[format]andtypes-jsonschemaas dev dependencies only.
Changed
-
Agentic scoring protocol
0.3.0.expected_policy_versionis derived from the delegation that covered the action instead of being copied from the attempt, and a policy-version-mismatch denial now records its covering chain. Every frozen Asteria Agentic v1 artifact is byte-identical under both protocols, because that world registers a single policy version; the number moves because the resolution rule changed, and a consumer scoring a world with more than one version would otherwise have no way to tell which rule produced their truth. -
agent-authority-contract/README.mdno longer states thatjsonschemawill become a project dependency. The validate command uses the pydantic models instead, because the schemas are generated from those models and the two are not nested — each accepts input the other refuses, so runtime schema validation would enforce a different surface than the scorer.
Added
- An AGENTIC_BENCHMARK.md “Trace conventions” section (issue #34) documenting the deterministic delegation-chain, evidence-reference, and side-effect rules the agentic evaluator grades against for the frozen Asteria Agentic v1 fixture, referenced from DATA_DICTIONARY.md.
0.10.0 - 2026-07-27
Added
- Relational integrity validation for custom agentic worlds, including reusable owner/runtime graph helpers, bounded v1 root and child delegator provenance, exact canonical-binding joins, and explicit integrity errors before evaluator truth generation.
- Agentic provenance exact-match and micro-precision metrics so fabricated evidence is distinguishable from missing evidence.
Changed
- Agentic reports now use scoring protocol
0.2.0; other task scoring versions and every frozen Asteria Agentic v1 artifact remain unchanged. - The README and Hugging Face dataset card now state the Python 3.12 minimum explicitly, and published copyright notices consistently name Redoubt Labs ltd.
- Package verification now derives the wheel filename from the project version,
so release bumps do not leave
make cichecking a stale distribution.
0.9.0 - 2026-07-27
Added
- Asteria Agentic v1 (issue #23): a frozen, checksum-bound procurement
conformance world with separate organisation, principal, logical-agent,
runtime, credential, resource, delegation, policy, and evidence roles; 24
replayable events and 11 positive/negative authority cases; physically
separate public and evaluator artifact trees; a nullable observed-action
JSONL contract; independent identity, authority, temporal, attribution,
owner, provenance, and side-effect metrics; public-only naive baselines; and
generate-agentic/evaluate agenticCLI support. - Reusable
synthworld.agenticcontracts, deterministic index/timestamp replay, field-by-field benchmark projection, tamper-checked package loaders, and an open-string case label so later custom worlds are not forced to reproduce Asteria’s canonical case set. - A repository-maintained Hugging Face dataset card that documents the Asteria public/evaluator boundary, authoritative artifact digests, and raw-file verification workflow.
Changed
- The extraction evaluator now rejects predictions that reference pages outside
the public corpus (for example, from a mismatched seed or persona count) with
EvaluationInputError, andExtractionPredictionSetrejects duplicate pages — consistent with the malformed-submission handling of the other scorers. - BENCHMARKS.md renders its three visuals as native Mermaid diagrams again,
dropping the committed
assets/*.svgfiles. GitHub renders Mermaid inline, so the document no longer depends on the image proxy. - The README now starts with a goal-led use-case chooser, and a new
USER_GUIDE.mdexplains the public-input-to-score workflow, current use cases, runnable commands, metric interpretation, and the safety boundary in plain language. - The all-task example now derives every prediction from public observations only and can write five CLI-ready prediction files, including an Asteria observed-action JSONL trace. The README, user guide, examples guide, and Asteria guide now document the complete export, integration, scoring, and metric-interpretation workflow. The roadmap use-case map now labels packaged, partial, and planned capabilities explicitly.
0.8.0 - 2026-07-22
Added
- Unified evaluation SDK (issue #1): versioned, oracle-free prediction schemas
and four scorers —
evaluate_extraction,evaluate_entity_resolution,evaluate_relationship_inference,evaluate_risk_calibration— that load truth themselves and return a uniformEvaluationReportof metrics and failure slices, with undefined metrics reported asnulland malformed submissions rejected viaEvaluationInputError. Metric definitions are versioned byscoring_version; the evaluation schemas are provisional0.1.0while the package is pre-1.0. Asynthworld evaluate <task>CLI scores a predictions file (with an optional--summarytable), andexamples/evaluate_all.pydemonstrates every task. - Separated exact-span extraction benchmark (issue #13): a product-safe
PublicExtractionCorpusand a physically separateExtractionAnswerKeyCorpus, joined and integrity-checked byExtractionBenchmark. Newgenerate-public-extractionandgenerate-extraction-answersCLI commands,generate_extraction_benchmark, and separately checksummedextraction-public-golden-v1.json/extraction-answer-golden-v1.jsonfrozen artifacts. The existing annotatedExtractionCorpusbundle is unchanged. - A DATA_DICTIONARY.md section for the extraction schema, distinguishing the annotated evaluator bundle from the product-safe projection.
- The frozen golden benchmarks are published as a browsable Hugging Face dataset (Bluntmachetti7/synthworld-benchmarks), linked from the README.
- BENCHMARKS.md (issue #11): naive baseline results over the extraction,
entity-resolution, relationship, and risk benchmarks, plus deterministic SVG
visual demonstrations (under
assets/) pulled straight from the pinned-seed corpora. Generated byexamples/generate_benchmarks_doc.py, whichmake baselineschecks for drift in CI.
Changed
- The extraction example now feeds the system under test only the public pages and loads the answer key separately to score.
0.7.0 - 2026-07-20
Added
- PyPI release workflow using GitHub OIDC Trusted Publishing, gated on the full CI suite and a tag-to-version match.
py.typedmarker so type checkers consume the package’s inline annotations; its presence in the wheel is asserted bymake package.examples/with a worked exact-span extraction evaluation and annotated sample output; the example runs as part ofmake ciso it cannot rot.- Project URLs, keywords, and classifiers in the packaging metadata.
- This changelog, a code of conduct, issue templates, and a pull-request template.
Changed
- README documents the
idcognito-synthworldinstall name, adds status badges, and links the examples. - The public data dictionary no longer references internal roadmap milestones.
0.6.0 - 2026-07-20
Initial public release, extracted from a private workspace with history squashed; internal 0.x iterations are not part of this repository.
Added
- Deterministic seeded world generator: personas, identity attributes, and
evidence-backed relationship ground truth (core-world schema
1.0.0). - Exposure corpus generator for breach, broker, search, and social scenarios
(exposure schema
1.0.0). - Exact-span extraction corpus with evaluator-only answer keys (extraction
schema
1.0.0). - Adversarial and relationship connection benchmarks with a strict
public/oracle type boundary (connection schema
1.0.0). - Risk-calibration benchmark with public observations physically separated
from evaluator-only score and factor truth (risk schema
1.0.0). synthworldCLI with eleven generate and metrics subcommands.- Seven frozen golden benchmarks with SHA256 manifests and byte-equality tests.
- Quality gates: strict mypy, ruff, 100% enforced branch coverage, an honesty gate for unexplained skips, CI on Python 3.12 and 3.14, and a full-history secret scan.