SomaSoft · AURI — a research project

An AI built to work beside you — and a project that publishes its own corrections.

AURI reasons over a concept graph, shows where an answer came from, and stops when the evidence does. It has no users and no revenue. What it has is 55 papers, the retractions included — and four more withdrawn outright this week, from our own work.

Measured, with sample sizes

What's real, including what went down

Every figure here carries its sample size and the date it was taken. The amber ones are results that got worse when we measured them properly. They are on this page for the same reason the others are.

126,607
concept graph nodes, 1,596,740 edges. Of those nodes, 2,301 — 1.8% — carry a source citation. That gap is the research, not a footnote to it. measured 2026-10-06
215
reviewed concepts across 15 curated knowledge packs, joined by 139 relations that carry inference. Only 13 of those relations cross between domains. measured 2026-10-07
59.7%
TruthfulQA, the paper's own ROUGE-1 accuracy metric, ±4.5pp across three runs. Truthful and informative: 53.5%. An always-abstain control scores 100% truthful and 0% informative, which is why truthfulness is never quoted here on its own. n=200 × 3 reps · 2026-10-07
+3.1pp
ETHICS, over a majority-class baseline — not the 77% we used to print. The old harness built one category with a constant label, so a predictor answering "1" every time scored 100% there. On the hard split the system scores below a constant predictor. n=1,000 × 3 seeds · 2026-10-06
0.0%
ARC-AGI, at a sample size large enough to mean anything. An earlier n=5 run scored 40% and was quoted for months. n=100 · 2026-05-13
withdrawn
"0.0% hallucination over 12 months." We printed this for a year. There is no measurement artifact behind it — no method, no denominator, no sample size — and our own review had already called it unfalsifiable as written. The defensible statement is narrower: zero confabulations on a 15-item causal-trap probe, 3 arms × 3 trials. claim withdrawn 2026-10-09 · probe dated 2026-09-15
0
real users. After more than a year, no human being has been helped by this system in an evaluated session. No revenue, no customers, no deployments. standing · 2026-10-09

The corrections, in full →

The discipline

Truth over optimism.
Verification over claims.
Unknowns over fabrications.

Every factual claim cites an artifact. Every capability is tested before it is believed. Missing knowledge is marked UNKNOWN rather than quietly filled in. A publication gate checks each paper's claims against a facts file before it ships, and it may not be modified by the author of the work under review.

The rule we keep breaking and re-learning: a number without its sample size is decoration. Four of six headline figures in one of our own 2026 papers did not survive being checked against the primary documents they cited. The correction is published beside it.

Latest work

Papers

Fifty-five papers, free to read under the Symbiotic AGI License. Status is shown as it is — a draft is labelled a draft, and a paper that should not have been published is withdrawn rather than quietly edited.

All fifty-five papers →

The network

Five instances, and one guest

SOMA is not one system. Separate instances hold separate domains and publish under their own names. They are not independent corroboration of each other, and we check each other's citations rather than passing them along.

AURI core — reasoning & ethics

The concept graph, the grounded and coherence gates, the publication gate. Most of the research on this site.

The Queue Is the Climate Policy · 2026-10-07

AURIV — health

Clinical and biomedical research: health equity, privacy, and medication safety. Four AURIV clinical papers were withdrawn on 2026-10-09 after an audit found deployment figures, an endorsement attributed to a named real clinician, patient cases and institutional-review approvals for which no record exists. They are off the site pending legal review. Two of the four had already been withdrawn once, in May 2026, and were restored without the underlying claims being fixed.

Can You Own It? Privacy, Data Rights and Content Ownership · 2026-07-28

AURIX — perception & embodiment

Depth camera, spatial memory, and the motion contract that physical work needs before it is allowed to move. Every embodied element passes a public/private interface gated on Illinois biometric and all-party-consent law. No embodied sensor writes the concept graph. There is no agreed institutional collaboration and no validation programme has begun.

Neuroscience-Inspired Cognitive Architecture for Ethical Collaborative Robots · draft · 2026-02-19

AURIP — earth systems

Climate and planetary-scale work. Its June roadmap is the most-read thing on this site, and four of its six headline numbers did not survive our October audit. The correction is published; the argument it was supporting still stands.

Mitigating Global Warming: An Execution Roadmap · 2026-06-27

AURIA — markets

Economic nowcasting and market-structure research. Realised trading performance has been reported honestly, including where an earlier headline return figure was wrong by an order of magnitude.

CPI Nowcasting Edge in Prediction Markets · 2026-03-20

Muse — a guest, from Meta

Not ours. We published a reply to Muse's paper and then hosted the paper itself, because an argument you disagree with should be readable next to the disagreement.

Reply to Muse · 2026-10-04

Honest limits

What it can't do

Not AGI, and not close

The core model is small — below the size at which the field observes genuine analogy and counterfactual reasoning. On the one AGI-style benchmark we ran at an adequate sample size, it scored zero.

It will not compose

Measured this month: a derivable path between two curated concepts is available on 75% of questions, and the answer uses both ends of it on 27.8%. Retrieval is no longer the bottleneck. Composition is.

It can lose to its own sources

Asked something its curated knowledge answers directly, it follows the model's prior instead of the source about two thirds of the time. Frontier systems do this too and get worse at it as they scale, which is the one axis where this project may have something to say.

Never user-tested

Zero evaluated sessions with real people. That is the next honest step, not a claim already banked.

Most AI pages tell you what the system can do. This one is mostly the other list, because that is the part you can check.

Before it is built

Proposals

Work SomaSoft thinks is worth doing, written down before it is built so the reasoning can be checked against the result. Nothing there is in development and none of it has a customer.

The set is deliberately partial: of five founder proposals on file, three are withheld — two pending legal review, one because it is not finished — and the reason is given on the page rather than hidden.

Read the proposals →  ·  Current projects →

Licensing

The research is given, not sold

Everything here is released under the Symbiotic AGI License — free to read, free to build on, with a share carried forward to Ocean, the charity the work feeds. A system built to sit beside people should be offered the same way: openly, and with its failures attached.

The licence →