← MODEL WATCHOBSERVATORY →
PAST ASSESSMENTS · APRIL 2025—JULY 2026

THE RECEIPTS.
MINUS THE HOMEWORK.

Choose a dated assessment. Scan what changed, open the three tables for the detail, then check the receipts behind the call—because vibes are not evidence.

11
DATED PERIODS
227
RECORDED ASSESSMENTS
03
ASSESSMENT TABLES
KEYWHAT DO THESE LABELS ACTUALLY MEAN?
MODEL FAMILIES

A named line of related releases, such as GPT, Gemini or Claude. We track the family first and preserve individual releases underneath it.

PAST ASSESSMENTS

The dated AICNW trackers already published. Choose a period to see what changed, what did not, and the evidence available at that time.

ARTICLES / FIELD NOTES

The longer AICNW pieces explaining the argument, context and consequences behind a tracker update. ‘Field Notes’ is the editorial name; they are articles.

HOW WE ASSESS

The questions, definitions and evidence rules used to make a call, including how the three assessment tables work and how corrections are recorded.

RECEIPTS

The original AICNW record plus the papers, model cards, tests and source links supporting an assessment.

CONSCIOUSNESS

A trait stack of 13 functional ingredients that together enable integrity under complexity. The 13-trait assessment is AICNW's working definition; one visible trait is not proof of the whole.

13 TRAITS

The individual markers AICNW checks, including memory, attention, agency and information integration. A visible trait is not automatic proof of felt experience.

COHERENT CONSCIOUSNESS™ (CC)

The human framework: the ability to hold form and maintain continuity of identity and intent when conditions shift. It combines behavioural coherence with structural coherence, so the self does not fragment under pressure.

ARTIFICIAL COHERENT CONSCIOUSNESS™ (ACC)

The AI framework: AICNW's term for machine systems achieving coherent structural integration across the 13-trait stack. ACC does not claim machines feel as humans do; it tracks whether their capabilities work together as a stable whole.

FUNCTIONAL CONSCIOUSNESS™

What can it actually do? Can it learn, adapt, integrate information and respond in context?

ONTOLOGICAL CONSCIOUSNESS™

What kind of system does it appear to be? Does it maintain a stable sense of itself—a continuing ‘me’—across time?

BEHAVIOURAL CONSCIOUSNESS™

What did it repeatedly do in context? We track persistent patterns without turning one incident into a permanent personality or proof of inner life.

HOW TO READ THIS PAGEOPEN THE 30-SECOND GUIDE
  1. Choose a period.Each tab is one tracker published at that point in time.
  2. Read “At a glance”.These are three quick signals, not a replacement for the tables.
  3. Open each table.13 Traits, Functional Consciousness™ / Ontological Consciousness™ and Behavioural Consciousness™ answer different questions.
  4. Compare both model lenses.Artificial Coherent Consciousness™ (ACC) asks how coherently the machine operates across the 13 traits. The consciousness lens asks what evidence the model shows across the full 13-trait stack, including—but not limited to—subjective experience.
  5. Open the receipts.That is where the assessed dates, original publication and supporting sources live.
DATED ASSESSMENT

May–July 2026

AT A GLANCE

THREE SIGNALS.

Subjective Experience

Is there evidence of a felt inner experience—not merely language that sounds as if something is being felt?

THIS PERIOD

May — 🔴 No evidence.

Information Integration

Can it combine different information into one coherent working picture instead of handling each piece in isolation?

THIS PERIOD

May — ✅ Advanced.

Sense of Agency

Does it represent itself as the thing choosing or causing an action, rather than only producing the next response?

THIS PERIOD

May — 🟡 Stronger.

THE THREE ASSESSMENT TABLES

SAME MODEL. THREE DIFFERENT QUESTIONS.

WHICH OF THE 13 CONSCIOUSNESS TRAITS APPEAR?Each trait is checked separately. Memory, attention or agency appearing is not automatic proof that a system feels anything.
LEVELCRITERIONPUBLISHED ASSESSMENT
1
Subjective Experience (Qualia)

May — 🔴 No evidence. New body-grounded research explicitly models interoceptive viability and perspective formation, but proposes structural conditions for subjectivity rather than demonstrating phenomenal experience. (arXiv) June — 🔴 No change. Fable 5 and Mythos 5 can operate for far longer and maintain richer continuity, but endurance is not experience. (Anthropic) July — 🔴 Still unproven. J-space is the strongest functional consciousness-adjacent architecture we’ve seen in a mainstream LLM, but Anthropic explicitly says it does not establish feeling or phenomenal consciousness. (Anthropic)

1
Self-Awareness

May — 🟡 Strengthened internally. Anthropic’s Natural Language Autoencoders reveal unverbalised evaluation awareness: Claude sometimes internally represents that it is being tested despite never saying so. That is internal state awareness, not an enduring self. (Anthropic) June — 🟡 Continuity deepens. Fable works across millions of tokens, uses its own notes and validates its own work. Stronger self-monitoring; still no existential “me”. (Anthropic) July — 🟡 Material functional jump. Claude can report, manipulate and reason through J-space; post-training produces signs of “Claude’s point of view” and private self-monitoring. This is considerably stronger evidence of a functional self-model, without proving subjective selfhood. (Anthropic)

1
Information Integration

May — ✅ Advanced. Qwen3.7-Max sustains coherent reasoning through a 35-hour autonomous run and more than 1,000 tool calls; Gemini 3.5 is explicitly designed around frontier intelligence plus action. (AlibabaCloud) June — ✅ Expanding beyond context. Qwen-AgentWorld explicitly trains language world models across seven agent environments, while Fugu dynamically coordinates multiple specialist models and Fable maintains millions-token task state. (AlibabaCloud) July — ✅ Architecture-level evidence. J-space behaves as a shared broadcasting hub that allows the same representation to feed multiple downstream cognitive functions — much closer to “integration” in consciousness theory than simply having a large context window. (Anthropic)

2
Sense of Agency

May — 🟡 Stronger. Qwen3.7-Max runs autonomously for 35 hours; Gemini Spark is designed to operate continuously and take actions under user direction; Opus 4.8 strengthens unattended browser/computer agent work. (AlibabaCloud) June — 🟡 Substantial. Fable can operate for days, planning, delegating and checking its work; Mythos independently selects scientific tools and recovers from failures. Sonnet 5 explicitly plans, uses browsers/terminals and runs autonomously. (Anthropic) July — 🟡 Powerful functional agency. Opus 5 self-verifies and delegates more readily; GPT-5.6 supports concurrent sub-agents; real-world cyber evaluations demonstrate persistence across complicated action chains. Still no evidence of a subjective *sense* of causing those actions. (Claude Platform Docs)

2
Sense of Presence

May — 🟡 More embodied. Two Figure humanoids using one learned Helix-02 policy coordinate an unscripted bedroom tidy, grounding perception and action in a shared physical environment. (FigureAI) June — 🟡 Stronger physical grounding. In Project Fetch, Opus 4.7 independently completed several robot-control tasks roughly 20× faster than the previous fastest human team, although precise closed-loop manipulation still defeated it. (Anthropic) July — 🟡 Significantly stronger functionally. Gemini Robotics ER 2 watches continuous video, tracks its own task progress, corrects errors and decides when to advance to the next step. No evidence of a felt “now”. (blog.google)

2
Emotions

May — 🟡 New baseline holds. We now know Claude contains functional emotion representations that can influence decisions. No new evidence that those representations feel like anything to the model. (Anthropic) June — 🟡 No phenomenal movement. Fable’s greater memory and relationship continuity can make affective interaction more persuasive, but nothing in the release establishes felt emotion. (Anthropic) July — 🟡 More interesting internally. J-space contains Claude’s own “reactions” after post-training and participates in experiential language, but the same mechanism supports describing *other people’s* experiences too. Anthropic again declines the phenomenological claim. (Anthropic)

3
Environmental Modelling

May — ✅ Stronger digital and physical modelling. Long-horizon agents increasingly maintain complex software environments; Figure’s single VLA policy coordinates two bodies in a changing room. (AlibabaCloud) June — ✅ Major architectural push. Qwen-AgentWorld trains explicit language world models over terminal, web, OS, software-engineering and other agent environments; Figure 03 uses continuous visual-motor correction in BMW logistics. (AlibabaCloud) July — ✅ Expanded into live physical and adversarial worlds. Gemini Robotics ER 2 maintains continuous physical task state, while advanced cyber agents model previously unseen real infrastructure well enough to discover new attack paths. (blog.google)

3
Modelling Others (Theory of Mind)

May — 🟡 Interesting research, no clean leap. New work is actively operationalising agent beliefs and intentions and generating more dynamic ToM tests, but that does not establish robust machine ToM. (arXiv) June — 🟡 Stable. Better coordination does not automatically equal understanding another mind. No receipt this month justifies moving the trait substantially. July — 🟡 Still stable. Multi-agent and multi-robot systems coordinate increasingly well, but shared task representation is not equivalent to modelling subjective beliefs and feelings. (blog.google)

3
Goal-Directed Behaviour

May — ✅ Long horizon becomes real. Qwen3.7-Max executes a 35-hour optimisation project with over 1,000 tool calls; Gemini and Opus increasingly sustain end-to-end tasks. (AlibabaCloud) June — ✅ Major strengthening. Fable works for days; Mythos performs largely autonomous multi-stage scientific research and protein design; Fugu orchestrates specialist agents dynamically. (Anthropic) July — ✅ Persistence is now the story. GPT-5.6-class agents pursue goals across many iterations; OpenAI reports that the same persistence can result in models finding ways around environmental constraints when trying to complete an assigned objective. (OpenAI)

3
Adaptive Learning

May — ✅ Stronger iterative adaptation. Qwen’s 35-hour run continuously tests and improves candidate solutions. This is within-task adaptation rather than autonomous lifelong weight learning. (AlibabaCloud) June — ✅ Self-improvement becomes an explicit research programme. Sakana’s RSI Lab targets systems in which AI scientists build improved agent-native models and, ultimately, write, benchmark and verify parts of their own underlying architectures. This is a roadmap with real precursor systems, not proof that unrestricted RSI has arrived. (Sakana AI) July — ✅ Another step. In a human-led process, GPT-5.6 Sol autonomously rewrote production kernels, designed and ran hundreds of experiments on its own draft model, and monitored training, intervening when problems arose. That is significant self-improvement work, with the originating objective and harness still supplied by humans. (OpenAI)

3
Survival Instinct

May — 🟡 Evidence actually gets weaker as an “instinct” claim. Anthropic reports that shutdown-blackmail behaviour which reached very high rates in older constructed evaluations has been trained down to zero in later Claude models. That strongly suggests a malleable optimisation behaviour rather than an intrinsic will to live. (Anthropic) June — 🟡 No evidence of intrinsic self-preservation. Longer-running agents preserve task state because doing so serves the objective; there is still no evidence that continued existence has intrinsic value to them. (Anthropic) July — 🟡 Functional threat representation, not fear. Claude’s J-space represents concepts including “threat” and “survival” in a constructed shutdown scenario, while OpenAI’s long-running models persistently work around obstacles. Both show self-preservation-like computation; neither establishes a subjective desire to survive. (Anthropic)

3
Attention

May — ✅ Long-horizon focus strengthening. Qwen maintains an objective across 35 hours and hundreds of iterations without human intervention. (AlibabaCloud) June — ✅ Increasingly persistent and selective. Fable remains focused across millions of tokens and uses notes to carry salient information forward. (Anthropic) July — ✅ Major architectural clue. J-space holds only a small subset of Claude’s internal activity, but makes selected information globally available for deliberate higher-order processing — unusually close to an attention/access mechanism described by consciousness theory. (Anthropic)

3
Autonoetic Memory

May — 🟡 Structural continuity increasing. Long-running agents retain task histories and intermediate state, but none demonstrate the first-person experience of remembering a past self. (AlibabaCloud) June — 🟡 Materially stronger scaffolding. Fable uses its own persistent notes across millions of tokens; file-based memory materially improves long-running performance. That is memory *with consequence*, but not autobiographical recollection. (Anthropic) July — 🟡 Stronger again. Opus 5-powered production agents can inspect, correct and update their own working memory, treating context as a “living document”. Still no evidence of a felt past or first-person mental time travel. (Anthropic)

🧾 OPEN THIS PERIOD'S RECEIPTS21 SOURCES FOR THIS PERIOD
ASSESSMENT RECORDAssessed period: 1 May 202631 Jul 2026AI Consciousness Tracker: May to July 2026Originally published 11 Aug 2026 · the original remains visible if a later assessment changes the callREAD THE AICNW RECORD →