Independent · Single-operator · Est. 2025


CIRWEL Research

— The research program, including its failures

What I am testing.
What already failed.

This page records the questions under test, the experiments that have closed, and the results that came back negative or inconclusive. An experiment that failed stays on the page with the number that failed it.

I am one person. Everything below was set, run, and recorded by me against a fleet I operate myself.


— The receipts

Program
Longitudinal runtime-state measurement for AI agents · begun September 2025
Operator
Kenny Wang · sole operator, no staff · ORCID 0009-0006-7544-2374
Evidence base
CIRWEL's own production fleet · continuous operation since November 2025 · self-traffic, not adoption
Papers & data
UNITARES · Trajectory Identity · Digital Proprioception · Accountability Without a Trusted Center · datasets & models

§01 — The questions under test

Two questions, at two scales.

One asks what can be read from a single agent while it works. The other asks what can be held to account between principals who do not trust each other. They are separate questions with separate evidence, and only the first has results.

Question I · one agent

Does an agent's runtime state carry information its outputs do not?

Here, state means an operational estimate built from recorded behavioral and system signals. It is not privileged access to a model's hidden internal state, thoughts, or chain of reasoning.

UNITARES updates that estimate while an agent works and compares it with an operating reference. The research program asks whether the reading is worth taking at all.

1. Discrimination
Do recorded state patterns distinguish agents or workload classes, and do those differences persist over time?
2. Prediction
Does prior state anticipate a bad outcome better than simpler baselines? This is still open; an earlier negative-looking answer was withdrawn because the cohort and permutation structure did not support that inference.
3. Intervention
Does acting on the state estimate improve anything measurable? This remains untested at the required scale.

Question II · many principals

Can accountability hold between principals who share no root of trust?

Agents increasingly act across infrastructure built and operated by different principals who do not trust one another and share no central controller. The safety-relevant questions there are relational: which process acted, who authorized it, what delegation chain led to it, and which control boundary committed the external effect.

The proposed answer is per-principal governance mediated by verifiable attestation rather than a trusted center, evaluated against four control regimes: prompt-only, log-only, federated per-principal runtime governance, and a centralized-governor reference. The point of the fourth is to measure the price of decentralization rather than assume it.

Status · designed and frozen, not run
The evaluation plan was pre-registered and frozen on 2026-08-05, before any evaluated scenario set existed. No benchmark run has been executed against it, so nothing on this page reports a federated result.
What exists today
The federation primitives were exercised live against the deployed server on 2026-06-30: per-principal credential self-proof, cross-principal attestation, credential isolation, and lineage held as claimed rather than trusted. Both principals ran from a single host, and the server detected and labelled that co-location. That is mechanism validation, not two organizations with conflicting interests. The multi-host case is not built.
Preprint
Accountability Without a Trusted Center · design and pre-registered evaluation plan · DOI 10.5281/zenodo.21930161

§02 — Closed, including negatives

What the current record actually supports.

Everything in this section belongs to Question I. Question II has produced no results to record.

Result · calibration changes the reading

In one offline counterfactual replay over 13,310 production observations, 28.9% of risk-band assignments would change under the alternative class-conditioned reading.

That shows the calibration choice materially changes this instrument. It does not establish predictive lift or intervention benefit, and the exact disagreement rate moves with the snapshot and comparison definition.

Negative result · reconstructed per-agent reference did not beat persistence

Zero of seven streams beat last-value persistence in the registered read.

The result applies to a cold-started reconstruction rather than the correctly warmed deployed reference; three nominal agents are one Raspberry Pi; and the pooled cross-dimension binomial is anti-conservative. The honest conclusion is that this reconstruction did not earn its description, not that every possible warmed reference is refuted.

Method finding · organic traffic cannot answer every question

For one prediction analysis, the population of tool failures and the population carrying suitable state readings were disjoint. The join was empty. That makes the intended inference structurally unavailable on that traffic, not merely underpowered.

§03 — Open

Unresolved is not the same as negative.

The prediction question had an answer that read negative, and that answer was withdrawn because the analyzed cohort and permutation blocks did not support the claimed inference. The result remains provenance, not standing evidence for or against predictive lift.

A pre-registered confirmatory read is scheduled for December 2026 with a kill criterion written before the data is seen. Until then the bounded question is closed to ad-hoc re-runs.

The self-predictability work is also under-powered: roughly four effective agents is not a population on which to claim generality.

Some of what I set out to demonstrate, I cannot yet demonstrate.

§04 — How the program is run

Commitments that can cost a result.

Pre-registration with a kill criterion
Bounded hypotheses get a fixed read and written stop condition before the data is seen.
No convenient re-runs
A bounded question does not get repeated until the answer improves.
Claims carry scope
Operational observations, benchmark results, non-detections, and untested questions are kept separate rather than flattened into one confidence label.
Single-operator disclosure
Every deployment number here comes from a fleet I run myself. Event volume is self-traffic, not adoption.

§05 — What the program needs

A skeptic, not an endorsement.

The main limitation is that the same operator built the system, runs the fleet, and produced the current evidence. The most useful next test is independent or adversarial evaluation on data I did not generate.

If you work on evaluations, multi-agent systems, or runtime oversight and want to test it, the code and data are open. Get in touch.