MULTIMODAL
Text, images, sound, video and action walk into the same model. Integration is the interesting bit.
WHAT DOES IT ACTUALLY MEAN?
A multimodal model can receive or produce more than one kind of information: text, images, audio, video, spatial data or actions. Some systems translate each mode into a shared internal representation; others connect specialist components.
WHY ARE WE TRACKING IT?
Conscious experience is deeply integrated across senses. Multimodality therefore matters to the investigation—but accepting several file types is not evidence of a unified inner world.
KEEP ONE EYEBROW UP.
Do not confuse a model that can inspect an image and transcribe audio with a model that binds those signals into one stable perspective.
RECEIPTS & REVISIONS.
Use this definition as the reading key for Model Watch. A capability claim only moves an AICNW assessment when the dated evidence meets the framework—not because a model said something dramatic on the internet.