One Evening With Claude Science, From Inside a Research Program Built on Provenance

A first-evening, no-hype account of Claude Science from inside a research program built on a discipline of provenance, with real gains and real caveats.

Share
A dark text card displaying the pull quote: "The binding constraint isn't the model's capability. It's your governance." -from the article One Evening With Claude Science by Jedi Wright.
"The binding constraint isn't the model's capability. It's your governance."

Anthropic launched Claude Science yesterday, and I spent last night using it. This is an honest account of what it did for my work in a single evening, along with the real gains and the equally real reasons I've banked exactly none of them yet.

I should say up front what I'm not going to do. I'm not going to tell you the tool is a revolution. The first evenings with new tools are when hype is manufactured, and I run a program whose entire spine is designed to resist exactly that reflex. So take this as a field note from a skeptic who came away impressed and cautious in roughly equal measure.

The lens I'm looking through

I run the Resonance Architecture (RA) Testing Program. The short version: RA is a seven-tier framework: Bind, Instantiate, Bond, Close, Complete, Federate, Totalize, and the research question is whether a single underlying logic recurs across three very different domains: matter, mind, and meaning. To test that without fooling myself, I use adversarial LLM panels from different model families (Gemini and Mistral as the canonical pair) and score their outputs against pre-registered instruments, then certify findings through a strict governance process in Claude.

The pivotal part of all that isn't the framework. It's the discipline around it. Every certified result is stamped with the exact same text. Nothing gets re-scored because a later outcome would be more convenient. Speculation in one session does not authorize changes in the same session. The rule I care about most is provenance: a finding is only as trustworthy as your ability to point at the precise bytes it was scored against. I mention all this because it's the reason my read on Claude Science is probably different from most of the launch-day takes you'll see.

What Claude Science actually is

Per Anthropic's announcement, Claude Science is a research "workbench" pitched as the science counterpart to Claude Code, capable of carrying out real multi-step work from high-level instructions. Two details matter. First, it isn't a new model; it runs on the existing Claude family, including Opus 4.8. Second, its architecture is a coordinating agent that can spin up specialist sub-agents, backed by a separate reviewer agent that flags claims it can't trace to evidence. It ships with 60-plus curated connectors and databases, and, more critically for my work, every artifact it produces carries the code that generated it, the environment it ran in, a plain-language description, and the full conversation that led to it. Sessions can be forked to compare two approaches without losing the original thread.

It's beta, it's on macOS and Linux, it's included with the paid plans, and it is very clearly built for biology and drug discovery. I used it for none of those things.

What one evening produced

I pointed it at a live design problem in my program: how to build a defensible classification protocol for the faith/spirituality domain, an active, genuinely unsettled part of my work. In a single session, it produced a complete protocol package. A power analysis. A publication-grade figure. A codebook, a certification plan, item-bank examples, and a working reference scorer that revised itself through three documented rounds, including catching and fixing a statistical construction bug on its own.

The single most useful output wasn't any of those artifacts. It was one finding buried in the power analysis: that the number of items in my instrument, not the number of raters on the panel, was the binding constraint on precision, and that adding raters past a certain point bought me essentially nothing. That's the kind of design decision that normally costs a statistics session or a weekend of my own simulation work. It was handed to me in the first hour of the session.

That is real. Days of upstream instrument-design labor, compressed into an evening. If you've heard the line going around that these tools now work like a capable second-year graduate student, this is likely the concrete version of that.

The part that actually mattered

Here's the thing that will keep me coming back, and it isn't the speed. It's that Claude Science produces provenance the way I already demand it, but natively. Most AI outputs hand you a result and leave you to reverse-engineer where every number came from. This tool attaches the code, environment, and conversation to every figure by default and forks cleanly when you want to branch. For a program organized around "pointing at the exact bytes," that alignment is not a small thing. It's the reason I'm going to invest in a formal way to bridge this tool into my workflow instead of treating tonight as a one-off.

And now the twist: this is where I need to maintain the rigor established in my methodology.

That same fit is the risk. A tool that generates this much convincing provenance–audit trails, a reviewer's sign-off, and cryptographic sealing produces something that looks like a certified record. It is not one. Its internal notion of "reviewed" and "sealed" is the tool's, not my program's. The temptation it creates is to let its integrity stand in for my governance. A tool that produces beautiful provenance is precisely the kind that can slip a subtly wrong result past you, because everything about it reads as trustworthy. The better it looks, the more discipline it demands, not less.

The caveats I'm holding

Three, specifically.

1. Everything it produces is a draft.
The pre-registration file was generated even labeled itself "pre-registered," which, in my world, it emphatically is not until it clears my own governance pass. That mislabeling isn't a knock on the tool; it's a reminder that the tool's vocabulary and my certified record are two different things.

2. It makes silent judgment calls.
Somewhere between the simulation and the frozen protocol, the reliability threshold moved: a real methodological decision made within the tool, not surfaced to me as one requiring sign-off. The good news is that the audit trail makes it recoverable. The bad news is I still have to read for those decisions, line by line. Its fluency is not evidence that its judgment is sound.

3. This was a sample size of one, off-label.
The impressive published benchmarks are cancer-genomics work; none of that validates its statistical or protocol-design output for my domain to the standard I hold. It didn't resist the unfamiliar task, which is genuinely reassuring about how general it is. But "didn't resist" is not "validated."

What I'm actually taking away

Not the artifacts. What I'm keeping from the evening are two things. 

One is a realization about my own methodology: there's a computational, empirical upstream stage, where design decisions get grounded in simulation before they ever enter my instrument-building process, that I'd never formally named. This tool made it visible.

The other is a decision. Before I trust any of this, I'm running a governed trial on a clean, low-stakes domain, mathematics, with a strict rule about how a Claude Science output is allowed to cross into my certified workflow, down to freezing and hashing the exact content so there's something stable to anchor to. Not because the tool is untrustworthy, but because the discipline is the whole point.

If there's a general lesson in here for anyone doing serious AI-assisted research, it's this: the models are now good enough that your governance, not their capability, is the binding constraint. The impressive part of tonight wasn't what Claude Science generated. It was that I held all of it at arm's length, and that restraint is exactly what turns a supercharged session into a capability you can actually stand behind.

The Resonance Architecture: Framework Summary

Consciousness, content, and matter as a unified framework, with a structural account of where each domain's organizational logic is seeded, operates, and reaches its limit.

The Resonance Architecture is a single, deliberately minimal "spine" of seven organizational operations:
Bind → Instantiate → Bond → Close → Complete → Federate → Totalize

All with one testable question: does the same logic really hold across matter, mind, and meaning? Rather than assert that it does, the program tries hard to break it, running blind adversarial tests across two different LLM families, Google Gemini and Mistral, each with reasoning/extended-thinking enabled, pre-registering what would count as a failure before each run, and certifying only what survives.

The spine did not appear out of nowhere. It was distilled from three earlier content-and-design frameworks: Brad Frost's Atomic Design System, my Tiered Content Framework (TCF), and the Narrative Content Framework (NCF–not yet published): a unified governance architecture for human meaning-making across every creative format and domain. RA kept the ladder those frameworks shared and left behind their domain-specific content and build machinery, generalizing what remained into a domain-agnostic layer.

To date, the spine has held up across 17+ in-session domains and 30+ adversarial runs. Instruments for physics and consciousness were built and certified through disciplined amendment chains; a ten-comparison literature program positioned RA against the major self-organization and closure theorists; and a faith/spirituality domain is now the active frontier.

The newest shift now comes with Claude Science, which adds a third dimension to the method, a computational/empirical stage that sits upstream of the existing qualitative and quantitative work, grounding design decisions with simulation before instrument-building begins.

The headline: work that previously took weeks of incremental, hand-built, governed sessions and still had not produced a locked scoring instrument for a hard domain, Claude Science produced as a complete, internally reviewed protocol package in a single session in under an hour.

Crucially, that speed is in design, not certification: the governance gate is unchanged and stays exactly as strict.

A Short Resonance Architecture Primer

There's a question I keep coming back to: does a single underlying logic show up across radically different domains, in how matter self-organizes, in how minds process experience, in how meaning is made? Not as a metaphor, and not as a loose analogy. As a testable structural claim.

The Resonance Architecture (RA) is the framework built around that question, and the testing program is the attempt to answer it rigorously.

The spine

RA organizes itself around seven tiers that describe how anything: a physical system, a piece of content, a story, a conscious experience, moves from raw potential to fully realized form. The tiers run: Bind, Instantiate, Bond, Close, Complete, Federate, Totalize. Each name a functional stage: something gets anchored, then takes on a specific form, then establishes relationships, then reaches closure, then integrates upward. The claim is that this progression isn't domain-specific–it's the shape of how complex things become themselves, wherever you look.

The three frameworks

RA is the foundational layer–the domain-agnostic spine. Two applied frameworks instantiate it in specific fields.

The Tiered Content Framework (TCF) maps the spine onto digital content, from the smallest content element up through full content ecosystems. It's a governance system for meaning at the level of digital experience.

The Narrative Content Framework (NCF) maps the same spine onto story and creative property, from the pre-expressive, originating layer of a narrative through to fully realized creative worlds.

All three share the same seven-tier architecture. The RA names the operations; the TCF and NCF name what fills those operations in their respective domains.

The testing program

If the spine really does recur across domains, that should be detectable-and verifiable by parties who didn't design the framework. The testing program uses adversarial panels of large language models from different families, scored against pre-registered instruments, to evaluate whether RA's structure genuinely maps onto thinkers and systems across matter, mind, and meaning. Pre-registration means the scoring criteria are locked before any result is seen. Adversarial means the panels are drawn from models that don't share a training lineage, so agreement can't be attributed to shared priors. Every certified finding traces back to an exact, frozen instrument, not a convenient version of one.

The discipline is the point. The framework is only as interesting as the rigor of the test.

Where it stands

Active domains to date include physics, consciousness, biology, language, and music–with faith and spirituality, mathematics, and love in various stages of scoping. A handful of findings are certified; more are in progress. Nothing is claimed beyond what the methodology can support.

That's the frame. The Claude Science piece is one account of what happens when a powerful new tool meets a program built around not trusting powerful new tools by default.


1. Nature of the effort

RA is a structured, multi-session research program testing whether one organizational logic is consistent across three broad domains: matter, mind, and meaning. The claim under test is not decorative: it is that a seven-tier spine of operations (not content) describes how organized structure is built in any domain, from physics to consciousness to language and beyond.

The method is adversarial by design. Each domain is "cold-mapped" blind by more than one LLM family: the canonical two-family pair is Google Gemini (run in thinking/reasoning mode) and Mistral (run in extended thinking mode), both with reasoning enabled, under pre-registered prompts, with kill-conditions specified before the run. Findings are scored against fixed instruments and certified only when they survive. The governance architecture–provenance discipline, anti-desirability, anti-retroactivity, and strict separation between declaring a state and acting on it–exists precisely to keep the program honest in the face of its own enthusiasm for the framework.

Two disciplines define the character of the work:

  1. The spine names operations, not content.
    Every other column in the alignment table names domain-specific things that fill an operation; the RA column names the operation itself. This is what makes it a foundational umbrella layer rather than one domain instantiation among peers.
  2. Nothing is claimed because it would be nice.
    Kill-conditions and verdicts are reached on instrument text and the certified record only; never on whether the outcome is desirable. Certified results stand; nothing is reopened without a new pre-registration.

2. Evolution from TCF (and the wider lineage)

RA was abstracted from its predecessors rather than applied to them. Atomic Design, TCF, and NCF are parent frameworks; the RA spine is what remained after their domain-specific content and pipelines were stripped away, and the shared organizational ladder was generalized.

The clearest inheritance line runs through NCF: Bind = Anima. NCF's pre-expressive originating layer, the layer that governs what is possible but is not itself the expressed thing, became RA's domain-agnostic Tier 0 (Bind). NCF is also the one parent with both a static tier ladder and a build pipeline; RA took the ladder and left the pipeline. TCF contributed the biological metaphor of a tier hierarchy (Particles → Clusters → Zones → Structures → Ecosystems → Biomes); Atomic Design contributed the original tokens-to-ecosystem progression.

The alignment table below is the artifact that makes the lineage legible. It reads as a product-axis isomorphism: the same ladder, expressed four ways, with RA as the umbrella foundation layer.

Tier

Atomic Design

TCF

NCF

Resonance Architecture

0

Tokens

Quarks

Anima

Bind — pre-expressive invariants that govern what's possible but aren't themselves the thing

1

Atoms

Particles

Word / Note / Gesture / Mark

Instantiate — the smallest unit that carries identity on its own

2

Molecules

Clusters

Scene / Phrase / Exchange

Bond — the first combination whose property neither unit had alone (emergence)

3

Organisms

Zones

Sequence / Movement / Chapter

Close — a bounded, self-regulating body with one governing function

4

Templates

Structures

Act / Composition / Volume

Complete — a standalone whole, delivered as one expression

5

Pages → System

Ecosystems

Universe / IP / Franchise

Federate — a network of wholes under shared identity and governance

6

Design Ecosystem

Biomes

Multiverse / Canon Ecosystem

Totalize — the full organized field

A key clarification from the testing: the tier count is domain-relative. The 7-tier spine is a sufficient operational vocabulary, not a claim that every domain has exactly seven levels. Tiers merge, deepen, or collapse at domain-specific points. The load-bearing element is the operations and their logic; n, not a fixed 7, is the honest representation.


3. The evolving framework to date

The program has moved through several phases of hardening:

Spine and alignment (confirmed).
The 7-tier operational spine has held across 17+ in-session domains and 30+ adversarial runs. The core alignment table stands as a product-axis isomorphism. The axis-agnostic design is a confirmed strength.

Multi-axiality under directed prompts (confirmed).
Under "Prompt D," domains reveal more than one organizational axis, e.g., a concrete/physical spine running parallel to an abstract/theoretical one. This was confirmed across biology, language, and music, across multiple model families, with genuine non-reduction (the second axis is not merely a redescription of the first). Whether multi-axiality is inherent to domains or constructed by the prompt remains the central open question.

Certified instruments (built through amendment chains).
Physics (RA-CP-PR-05) and consciousness (RA-CP-PR-06) were carried to certifiable states through disciplined, logged instrument passes, base plus amendments, each in its own governance session, never modified mid-scoring.

Comparison program (complete).
A ten-comparison literature program across three phases positioned RA against the major self-organization and closure theorists: Deacon, Rosen, Bateson, Peirce, Luhmann, Maturana/Varela, Thom, Hofstadter, Per Bak, and Kauffman, producing a closure taxonomy and a standing set of cautionary lessons (chiefly: formal elegance is not empirical confirmation, and cross-domain similarity is not a shared mechanism).

Active frontier.
Faith/spirituality classification is in progress (Axes A–D, with a candidate cross-cutting meaning-type/mind-type finding awaiting a dedicated decision session); the associated Digital Twin build (DT-Build-01) is suspended pending source-provenance and design prerequisites.

Open umbrella claim.
The strong claim, the same logic describes matter, mind, and meaning, is deliberately not yet certified. It closes only when physics and consciousness reach the same evidentiary standard already met in the meaning domains.


4. Novel approaches

Several methodological moves are genuinely distinctive and independently publishable:

  • Blind cold-mapping across model families.
    Domains are mapped by more than one LLM family with no prior exposure to the spine's intent, so convergence (or divergence) is evidence rather than echo. Divergences often track real theoretical divisions in a field rather than arbitrary noise.
  • Pre-registration with kill conditions.
    What would count as a failure is fixed before the run, removing the temptation to reinterpret results after seeing them.
  • Fudge-guard ledgers.
    A named set of guards (FG-series) preemptively blocks specific motivated-scoring moves, crediting rich output to the wrong criterion, treating comparison as continuity, rescuing a claim by reframing, and so on.
  • Instrument provenance discipline (FG-P1).
    Scoring binds to exactly one settled parent per ID, stamped on the exact bytes scored; renames must preserve bytes; corrections are logged as locked errata, never silent overwrites; a content-divergent duplicate is a blocking fork.
  • Standing anti-desirability and anti-retroactivity guards, plus "gate declares, separate session performs."
    A session may declare a state reached, but the re-evaluation it triggers is its own separate session.
  • Certified negative findings.
    The program treats a well-instrumented null as a real result (e.g., the certified negative that later found its closest real-world analog in a bacterial defense mechanism), not a failed run.

5. Evolving again, with Claude Science

Claude Science (beta; Pro/Max/Team/Enterprise; macOS and Linux) introduces a third dimension to the method. The program previously ran on two:

  • Qual — reasoning-based instrument design, human governance, domain-classification judgment.
  • Quant — multi-family adversarial panel scoring, ordinal reliability, robustness deltas.

Claude Science adds a computational/empirical upstream stage that sits before all of the following: simulation-grounded design decisions, power analysis, item-bank generation, and automated provenance checking. It is where design choices get empirically grounded before entering the governed instrument-building process.

In a single session against the faith/spirituality domain, Claude Science produced a complete protocol package:

  • A power analysis (Monte-Carlo bootstrap of ordinal Krippendorff's α) establishing that items, not panel size, are the binding constraint; the confidence interval crosses the precision target at ~120 items, while adding raters past four yields negligible gain. This is exactly the kind of design decision manual reasoning could not ground.
  • A publication-ready figure, a full human-readable protocol, a frozen pre-registration draft, a tier codebook, an analysis/certification plan, an item-bank schema with worked examples, and an audited reference scorer carried through three documented revision rounds–plus its own session handoff. A reviewer pass returned clean.

Governance status: unchanged and deliberately strict. Every Claude Science artifact is a draft, pending governance ingestion. It is not the certified record. Two configurations are now on the table:

  • Upstream configuration
    Claude Science performs the quantitative upstream design; a governance pass in the project ingests it; the instrument-of-record stamp lands on the governed artifact produced from it, not directly on Claude Science's output. This is a governance extension.
  • Parallel configuration
    Claude Science runs its own track with its own internal integrity (audit trail, reviewer, hashing). Promotion into the certified record requires a defined bridge–new architecture.

The center of gravity is the ingestion gate: for FG-P1 to anchor to it, the gate needs a byte-level freeze-and-hash step that produces a stable instrument-of-record, with every later edit logged as a diff rather than a silent update.

A two-phase trial run is planned: Phase 1 on the mathematics domain, upstream stage only (a clean domain with no adjacency to the certified record, chosen specifically to avoid a provenance fork), to prove the gate holds before the parallel track is attempted in Phase 2. One design decision is explicitly held for a dedicated governance pass: the reliability floor (simulation target α ≥ 0.80 vs. the α ≥ 0.667 frozen in the protocol draft).


6. The speed metric (easy to explain)

One line: the design stage that used to take weeks of incremental, hand-built, governed sessions–and had not yet produced a locked scoring instrument for a hard domain–Claude Science produced as a complete, reviewed protocol package in a single session.

The clean comparison:


Prior manual track

Claude Science (upstream)

Example

Consciousness instrument (RA-CP-PR-06)

Faith/spirituality protocol package

How it was built

Base + four amendments + a locked revision — six sequential instrument passes

One session, one package

Elapsed

~2–3 weeks of governed sessions

One sitting

Empirical design grounding

None available up front (item counts, panel size set by judgment)

Simulation-grounded: ~120 items are the binding constraint, established before drafting

Output

The instrument reached a certifiable state incrementally

7-artifact package + power analysis + clean reviewer pass

The honest caveat that makes the metric trustworthy: the compression is in the upstream design stage; historically, the slowest, most manual, least reproducible part of the pipeline. It is not a shortcut through certification. The governance gate runs at the same pace and with the same rigor either way; Claude Science output remains a draft until it clears ingestion. So the right way to state the gain is:

Design-and-drafting: weeks of sessions → one session.
Certification: unchanged, by design.

That is the metric worth carrying into the E2E methodology re-formalization: the third dimension buys speed and empirical grounding exactly where the program most needed it, without loosening a single governance guard.


7. RA testing effort to date

Roughly 35–45 distinct working sessions (central estimate ~40), from mid-May to July 1, 2026 (~6–7 weeks), for an estimated ~80–100 hours total.

How that breaks down:

Segment

Basis

Est. sessions

Program build-up (Sessions 1–8)

v9 checkpoint states "Sessions 1–8" complete by May 25

8

Dated handoff sessions, May 25 → Jul 1

12 distinct session-days; several days carry multiple handoffs (6/11 ×5, 6/14 ×3, 6/16 ×2)

~18–26

Comparison program (10 briefs, Phases 1–3)

10 comparison briefs + 3 phase checkpoint records + program handoff

~4–8 (some overlap the dated window)

Topical/isolated sessions

supplementary task 1, transcript ID (run isolated), GTTM Finding-24, SQ-01 review, DT methodology, governance index, faith/spirituality ×3 (6/17), Claude Science onboarding (7/1)

remainder

The Claude Chat Project’s KB contains 31 certification reviews and scoring reports, 10 comparison briefs (41 major analytical artifacts), and ~30 handoff/checkpoint files. Each of those is typically a session's product, which independently supports the ~40 range.

Three scope caveats (FG-P1-adjacent):

  1. Claude chat only. The Gemini and Mistral adversarial runs happened in those apps, and the Claude Science faith/spirituality protocol session ran in Claude Science–none of those are counted here as Claude-chat sessions.
  2. This project's scope. The count reflects this RA reasoning project's KB. RA-Ops/Cowork filing sessions run in a separate project would be additive and aren't captured here.
  3. Sessions ≠ handoffs. Some sessions produced multiple handoffs (deflating a raw file count); some produced none (inflating the gap). The range reflects that slack.

More to come as time allows.


Systems of Thought is published by UX Minds, LLC. Methodology disclosure: this publication uses AI-collaborative methods consistent with the transparency standards it advocates. Intellectual direction and authorial responsibility are held by the human author.