How we build it
We do not describe the scene. We remember the space.
We read space as structure, remember the path travelled, speak differently to each person, and stay silent when we are not certain.
How this differs
A system that describes a scene and a system that models a space are not the same thing.
Image Frame · A single image
Single Frame
Current frame
Visual Recognition
Object recognition
Caption Generation
Sentence generation
Audio Output
Speech output
Single-frame inference · No spatial memory
Visual Input · Camera Stream
Spatial Parsing
Object · Position · Distance · Relation
3D Spatial Model
Scene Geometry / Spatial Graph
Persistent Spatial Memory
Object Tracking / Temporal Memory
Context & Relevance Engine
User State × Environment × Priority
Adaptive Description
What to say · When to say · How to say
↻ Update · The decision feeds back into the spatial model and the memory
Frame → Caption
Recognise the current frame and turn it into a sentence
VS.
Space → Model → Memory → Decision
Model the space, remember it, judge, and deliver only what is needed
Recognition coverage
It tells four kinds of situation apart.

Spatial tracking
Spatial Tracking
Walking along corridors and passages. People and walkable ground are given identifiers and followed as the same objects even as the view changes.
People and obstacles
People & Obstacle Detection
Crowded shops and markets. Standing people, passing people, display units, and furniture are each identified separately.
Scene understanding
Scene Understanding
Interiors full of glass walls and columns. Passable areas and blocked areas are separated surface by surface.
Navigation in complex space
Navigation in Complex Space
Alleyways packed with signs and stalls. We deliberately included the conditions that make recognition hardest.
Demonstration
This is it running.
A three-minute film of what we have actually built. For anyone who cannot watch it, the scenes are written out below.
What the film shows
- Opening
- What we build
- An introduction to the structure: turning what the glasses see into spatial information, and proposing the next action on top of it.
- 0:40
- First, decide whether to speak
- It looks first at the width available to pass and at nearby objects, then decides whether to speak at all. Most of the time it does not. When it does, it first composes a sentence of plain fact, then rewrites it in the tone and level of detail that suits the person.
- 1:00
- It does not lose track of people
- Moving things such as people and bags are given identifiers and followed as the same object even as the view changes. Filmed at a conference venue in Tokyo.
- 1:40
- We tried it with our eyes closed
- Moving through a real gallery with our eyes closed, relying on the glasses alone.
- 2:00
- Rebuilding the space from many videos
- Several recordings are combined to reconstruct the floor plan of the venue. Of 753 camera positions, 752 were registered.
Evidence
A record of the four running for real.
Spatial axis · World State
What is in front of me right now?
People, walkable ground, doorways, and glass walls are each identified as distinct things. Every one carries a mark showing whether the judgement is confirmed or merely inferred, so that nothing uncertain is ever spoken aloud.

Speech decision · Silence Policy
Should it speak right now?
There is always something to say: a person ahead, a doorway, a banner. We filter these through four stages. Has the passable width narrowed? Is this the first time seeing it? Is it dangerous? Was it just mentioned? Most candidates are filtered out here and end in silence.
Person axis · Persona
Does this person need to hear it?
From the same scene, three different sentences go out. For someone blind from birth, only what is needed to move, kept short. For someone who lost their sight later, the scene as it would have looked. For an older user, slower and warmer. It is not a change of tone; the choice of what to say diverges first.
Spatial reconstruction · Multi-video
What kind of space is this?
Five recordings, taken separately with glasses and with a drone, are merged so that every camera position lands in one coordinate frame. Of 753 positions, 752 were registered. From those points we fitted seven walls and eleven barrier posts into a plan, and shortlisted six likely locations for exhibited works.

Logs
Screens we kept while building.
Screen recordings we kept while building. Not a highlight reel — these are the runs as they came out that day.
These are logs from before human review. As the warning banner on the footage says, we do not use the numbers on these screens as performance figures. The reviewed numbers are in the sections above.
What it measures while walking, and what it lets pass
What the glasses saw, and the decision made at that moment, on one screen.
What it measures each step, and the path behind it · 20 sec
The top is what the glasses saw. Walkable floor is painted green, each person gets a box, and depth appears at the top right. The middle is the decision: the passable width is wide enough, so say nothing. The bottom draws the path already walked, split into open space, corridor, and narrow corridor.
What the glasses piled up over 45 seconds · 19 sec
Even while saying nothing, the glasses keep looking. Text read and kinds of objects encountered accumulate, and the path walked is drawn on the right. Eyes only see the moment they are raised; the glasses see the whole walk. That difference is what the recap gives back later.
Even at a crowded entrance, it stays quiet · 1 min 40 sec
The passable width has narrowed by half. It still says nothing, because it spoke a moment ago. Having spoken recently outranks having narrowed. Having a reason to speak is not the same as having to speak.
Cutting things apart, and keeping their numbers
People and objects are separated one by one and keep the same number as they pass. No faces are used.
People and their bags each get a number · 2 min
People, handbags and chairs are cut into separate pieces and numbered. The same number follows a person as they overlap and walk past. The top left records how many were detected this frame, how many are visible, how many are new, and how many are ambiguous.
Confirmed and estimated are kept apart · 1 min 14 sec
The same scene is split into confirmed and estimated. What is not certain is neither counted nor spoken. Classes we have decided to stay silent about drop off the list entirely.
The same screen, in several places
On the left is what the glasses saw; on the right is the World State at that moment. What is there (objects), what kind of space it is, and what might be about to happen (behaviour candidates) are written separately. We ran the same screen in different places to see what breaks first when conditions change.
A carpeted corridor · 1 min
A corridor with several people passing through. A busy floor pattern makes the walkable surface harder to hold on to. The bottom right also explains what each on-screen marking means.
An indoor shopping street with bright signs · 1 min
Lit signage and a lot of glass. Text reads well here, but glass and reflections unsettle the walkable-floor judgement.
Dark surroundings, bright signs · 4 min 8 sec
Everything is dark except the signage. We ran this one long to see how depth estimation wavers where the brightness range is extreme.
A booth covered in text · 4 min 11 sec
A booth with a laptop and a display panel. Here a behaviour candidate — someone might linger in this spot — actually fires. It was the only one of the five places where that happened.
A crowded exhibition hall · 5 min
The busiest condition of the set. Five minutes without a cut, to see whether the screen holds while people keep moving in and out.
Rebuilding a space from several videos
How separately shot footage is brought into one coordinate system.
From points to a floor plan · 24 sec
Camera positions are aligned to build a point cloud, several videos are laid onto the same coordinates, and then walls, openings, artworks and routes are attached to make a plan. The result of this process, with its numbers, is in demonstration 04 above.
In the field
Where we filmed.
Scale
We hold agreements with three museums and two centres for blind people in Korea, and have filmed on site at more than fifteen indoor venues. Exhibition halls, conferences, shops, and streets were captured from a first-person view. We wrote our own specification covering thirty classes of spatial object, reviewed 2,037 frames sampled every second, then reviewed a further 6,822 frames at six-frame intervals.
Diversity
We filmed wearing five different models of AI glasses, and placed drone footage into the same coordinate frame. Bright exhibition halls, dark conference venues, crowded entrances, shops full of glass walls: we deliberately chose the conditions that make recognition harder.
Privacy
All processing finishes on the device. Recorded video is never uploaded to a server. We do not track faces or gaze, and we keep no per-person record of who looked at what, or for how long.
One memory, three uses
The same memory, read in three directions.
Experience Memory
Order travelled · Places lingered · Confirmed · Missed
While walking
Guide
Says only what is needed, in real time
After a trip
Recap
Retraces what one person confirmed and what they missed
After many visitors
Report
Tells the venue operator where people get stuck
These are not three separate products. They are one memory, read in three directions.
Where it runs
On the device, with no internet.
It runs on the device, with no internet connection. It works in underground passages and inside buildings where the signal is weak, and the recorded video never leaves the device.

Measured figures
Measured on real footage and real hardware.
Measured on 300 seconds of first-person footage filmed indoors, and on the laptop and edge device we actually use.
14 of 601
Moments where speech was possible in 300 seconds, and how many were actually spoken
0.13 sec
Time to read one scene. About the length of a blink
1.9 GB
Size of the model that sits on the device. Roughly a few hundred photographs
150 combinations
Unseen combinations of user conditions where instructions were still followed








