AI media workflows are starting to hit a different bottleneck: not generating or analysing individual assets, but keeping context consistent across an entire piece of content.
Iyuno is addressing that problem with CLOE, its contextual intelligence platform built around multiple specialised AI agents, multimodal processing and persistent memory.
The system separates visual, audio and language inputs before bringing them back together through what Iyuno calls Cognitive Integration. That layer combines multimodal fusion, reasoning, contextual analysis and memory formation, allowing CLOE to retain relationships, emotional intent, narrative progression and character continuity across a project.
Instead of treating every scene or task as a fresh input, CLOE creates a persistent Context Memory that can be reused across later workflows.
People don’t understand a story through dialogue alone or visuals alone. We naturally combine multiple sensory inputs, connect them with memory and reasoning, and build understanding over time. CLOE was designed to follow that same process. – Iyuno CEO David Lee

At the first layer, CLOE’s Human Sensory System uses separate agents to analyse elements including appearance, action, dialogue, tone, music and subtext.
Those signals are then combined into a shared contextual layer designed to support workflows across localization, accessibility, creative production and other media applications. Lee says:
Understanding isn’t created by a single model. It emerges when multiple perspectives are brought together into a shared contextual memory.
For Iyuno, the goal is to stop AI tools rebuilding context for every new task.
CLOE instead keeps dubbing, subtitling, accessibility and production workflows connected by the same evolving understanding of the story.
