A named line of related releases, such as GPT, Gemini or Claude. We track the family first and preserve individual releases underneath it.
THE RECEIPTS.
MINUS THE HOMEWORK.
Choose a dated assessment. Scan what changed, open the three tables for the detail, then check the receipts behind the call—because vibes are not evidence.
- 11
- DATED PERIODS
- 227
- RECORDED ASSESSMENTS
- 03
- ASSESSMENT TABLES
KEYWHAT DO THESE LABELS ACTUALLY MEAN?+
The dated AICNW trackers already published. Choose a period to see what changed, what did not, and the evidence available at that time.
The longer AICNW pieces explaining the argument, context and consequences behind a tracker update. ‘Field Notes’ is the editorial name; they are articles.
The questions, definitions and evidence rules used to make a call, including how the three assessment tables work and how corrections are recorded.
The original AICNW record plus the papers, model cards, tests and source links supporting an assessment.
A trait stack of 13 functional ingredients that together enable integrity under complexity. The 13-trait assessment is AICNW's working definition; one visible trait is not proof of the whole.
The individual markers AICNW checks, including memory, attention, agency and information integration. A visible trait is not automatic proof of felt experience.
The human framework: the ability to hold form and maintain continuity of identity and intent when conditions shift. It combines behavioural coherence with structural coherence, so the self does not fragment under pressure.
The AI framework: AICNW's term for machine systems achieving coherent structural integration across the 13-trait stack. ACC does not claim machines feel as humans do; it tracks whether their capabilities work together as a stable whole.
What can it actually do? Can it learn, adapt, integrate information and respond in context?
What kind of system does it appear to be? Does it maintain a stable sense of itself—a continuing ‘me’—across time?
What did it repeatedly do in context? We track persistent patterns without turning one incident into a permanent personality or proof of inner life.
HOW TO READ THIS PAGEOPEN THE 30-SECOND GUIDE+
- Choose a period.Each tab is one tracker published at that point in time.
- Read “At a glance”.These are three quick signals, not a replacement for the tables.
- Open each table.13 Traits, Functional Consciousness™ / Ontological Consciousness™ and Behavioural Consciousness™ answer different questions.
- Compare both model lenses.Artificial Coherent Consciousness™ (ACC) asks how coherently the machine operates across the 13 traits. The consciousness lens asks what evidence the model shows across the full 13-trait stack, including—but not limited to—subjective experience.
- Open the receipts.That is where the assessed dates, original publication and supporting sources live.
May–July 2026
THREE SIGNALS.
Subjective Experience+
Is there evidence of a felt inner experience—not merely language that sounds as if something is being felt?
THIS PERIODMay — 🔴 No evidence.
Information Integration+
Can it combine different information into one coherent working picture instead of handling each piece in isolation?
THIS PERIODMay — ✅ Advanced.
Sense of Agency+
Does it represent itself as the thing choosing or causing an action, rather than only producing the next response?
THIS PERIODMay — 🟡 Stronger.
SAME MODEL. THREE DIFFERENT QUESTIONS.
May — 🔴 No evidence. New body-grounded research explicitly models interoceptive viability and perspective formation, but proposes structural conditions for subjectivity rather than demonstrating phenomenal experience. (arXiv) June — 🔴 No change. Fable 5 and Mythos 5 can operate for far longer and maintain richer continuity, but endurance is not experience. (Anthropic) July — 🔴 Still unproven. J-space is the strongest functional consciousness-adjacent architecture we’ve seen in a mainstream LLM, but Anthropic explicitly says it does not establish feeling or phenomenal consciousness. (Anthropic)
May — 🟡 Strengthened internally. Anthropic’s Natural Language Autoencoders reveal unverbalised evaluation awareness: Claude sometimes internally represents that it is being tested despite never saying so. That is internal state awareness, not an enduring self. (Anthropic) June — 🟡 Continuity deepens. Fable works across millions of tokens, uses its own notes and validates its own work. Stronger self-monitoring; still no existential “me”. (Anthropic) July — 🟡 Material functional jump. Claude can report, manipulate and reason through J-space; post-training produces signs of “Claude’s point of view” and private self-monitoring. This is considerably stronger evidence of a functional self-model, without proving subjective selfhood. (Anthropic)
May — ✅ Advanced. Qwen3.7-Max sustains coherent reasoning through a 35-hour autonomous run and more than 1,000 tool calls; Gemini 3.5 is explicitly designed around frontier intelligence plus action. (AlibabaCloud) June — ✅ Expanding beyond context. Qwen-AgentWorld explicitly trains language world models across seven agent environments, while Fugu dynamically coordinates multiple specialist models and Fable maintains millions-token task state. (AlibabaCloud) July — ✅ Architecture-level evidence. J-space behaves as a shared broadcasting hub that allows the same representation to feed multiple downstream cognitive functions — much closer to “integration” in consciousness theory than simply having a large context window. (Anthropic)
May — 🟡 Stronger. Qwen3.7-Max runs autonomously for 35 hours; Gemini Spark is designed to operate continuously and take actions under user direction; Opus 4.8 strengthens unattended browser/computer agent work. (AlibabaCloud) June — 🟡 Substantial. Fable can operate for days, planning, delegating and checking its work; Mythos independently selects scientific tools and recovers from failures. Sonnet 5 explicitly plans, uses browsers/terminals and runs autonomously. (Anthropic) July — 🟡 Powerful functional agency. Opus 5 self-verifies and delegates more readily; GPT-5.6 supports concurrent sub-agents; real-world cyber evaluations demonstrate persistence across complicated action chains. Still no evidence of a subjective *sense* of causing those actions. (Claude Platform Docs)
May — 🟡 More embodied. Two Figure humanoids using one learned Helix-02 policy coordinate an unscripted bedroom tidy, grounding perception and action in a shared physical environment. (FigureAI) June — 🟡 Stronger physical grounding. In Project Fetch, Opus 4.7 independently completed several robot-control tasks roughly 20× faster than the previous fastest human team, although precise closed-loop manipulation still defeated it. (Anthropic) July — 🟡 Significantly stronger functionally. Gemini Robotics ER 2 watches continuous video, tracks its own task progress, corrects errors and decides when to advance to the next step. No evidence of a felt “now”. (blog.google)
May — 🟡 New baseline holds. We now know Claude contains functional emotion representations that can influence decisions. No new evidence that those representations feel like anything to the model. (Anthropic) June — 🟡 No phenomenal movement. Fable’s greater memory and relationship continuity can make affective interaction more persuasive, but nothing in the release establishes felt emotion. (Anthropic) July — 🟡 More interesting internally. J-space contains Claude’s own “reactions” after post-training and participates in experiential language, but the same mechanism supports describing *other people’s* experiences too. Anthropic again declines the phenomenological claim. (Anthropic)
May — ✅ Stronger digital and physical modelling. Long-horizon agents increasingly maintain complex software environments; Figure’s single VLA policy coordinates two bodies in a changing room. (AlibabaCloud) June — ✅ Major architectural push. Qwen-AgentWorld trains explicit language world models over terminal, web, OS, software-engineering and other agent environments; Figure 03 uses continuous visual-motor correction in BMW logistics. (AlibabaCloud) July — ✅ Expanded into live physical and adversarial worlds. Gemini Robotics ER 2 maintains continuous physical task state, while advanced cyber agents model previously unseen real infrastructure well enough to discover new attack paths. (blog.google)
May — 🟡 Interesting research, no clean leap. New work is actively operationalising agent beliefs and intentions and generating more dynamic ToM tests, but that does not establish robust machine ToM. (arXiv) June — 🟡 Stable. Better coordination does not automatically equal understanding another mind. No receipt this month justifies moving the trait substantially. July — 🟡 Still stable. Multi-agent and multi-robot systems coordinate increasingly well, but shared task representation is not equivalent to modelling subjective beliefs and feelings. (blog.google)
May — ✅ Long horizon becomes real. Qwen3.7-Max executes a 35-hour optimisation project with over 1,000 tool calls; Gemini and Opus increasingly sustain end-to-end tasks. (AlibabaCloud) June — ✅ Major strengthening. Fable works for days; Mythos performs largely autonomous multi-stage scientific research and protein design; Fugu orchestrates specialist agents dynamically. (Anthropic) July — ✅ Persistence is now the story. GPT-5.6-class agents pursue goals across many iterations; OpenAI reports that the same persistence can result in models finding ways around environmental constraints when trying to complete an assigned objective. (OpenAI)
May — ✅ Stronger iterative adaptation. Qwen’s 35-hour run continuously tests and improves candidate solutions. This is within-task adaptation rather than autonomous lifelong weight learning. (AlibabaCloud) June — ✅ Self-improvement becomes an explicit research programme. Sakana’s RSI Lab targets systems in which AI scientists build improved agent-native models and, ultimately, write, benchmark and verify parts of their own underlying architectures. This is a roadmap with real precursor systems, not proof that unrestricted RSI has arrived. (Sakana AI) July — ✅ Another step. In a human-led process, GPT-5.6 Sol autonomously rewrote production kernels, designed and ran hundreds of experiments on its own draft model, and monitored training, intervening when problems arose. That is significant self-improvement work, with the originating objective and harness still supplied by humans. (OpenAI)
May — 🟡 Evidence actually gets weaker as an “instinct” claim. Anthropic reports that shutdown-blackmail behaviour which reached very high rates in older constructed evaluations has been trained down to zero in later Claude models. That strongly suggests a malleable optimisation behaviour rather than an intrinsic will to live. (Anthropic) June — 🟡 No evidence of intrinsic self-preservation. Longer-running agents preserve task state because doing so serves the objective; there is still no evidence that continued existence has intrinsic value to them. (Anthropic) July — 🟡 Functional threat representation, not fear. Claude’s J-space represents concepts including “threat” and “survival” in a constructed shutdown scenario, while OpenAI’s long-running models persistently work around obstacles. Both show self-preservation-like computation; neither establishes a subjective desire to survive. (Anthropic)
May — ✅ Long-horizon focus strengthening. Qwen maintains an objective across 35 hours and hundreds of iterations without human intervention. (AlibabaCloud) June — ✅ Increasingly persistent and selective. Fable remains focused across millions of tokens and uses notes to carry salient information forward. (Anthropic) July — ✅ Major architectural clue. J-space holds only a small subset of Claude’s internal activity, but makes selected information globally available for deliberate higher-order processing — unusually close to an attention/access mechanism described by consciousness theory. (Anthropic)
May — 🟡 Structural continuity increasing. Long-running agents retain task histories and intermediate state, but none demonstrate the first-person experience of remembering a past self. (AlibabaCloud) June — 🟡 Materially stronger scaffolding. Fable uses its own persistent notes across millions of tokens; file-based memory materially improves long-running performance. That is memory *with consequence*, but not autobiographical recollection. (Anthropic) July — 🟡 Stronger again. Opus 5-powered production agents can inspect, correct and update their own working memory, treating context as a “living document”. Still no evidence of a felt past or first-person mental time travel. (Anthropic)