SightSpace

How we build it

We do not describe the scene. We remember the space.

We read space as structure, remember the path travelled, speak differently to each person, and stay silent when we are not certain.

How this differs

A system that describes a scene and a system that models a space are not the same thing.

ConventionalFrame-based Scene Description

Image Frame · A single image

  1. Single Frame

    Current frame

  2. Visual Recognition

    Object recognition

  3. Caption Generation

    Sentence generation

  4. Audio Output

    Speech output

Single-frame inference · No spatial memory

SightSpaceSpatial Intelligence Pipeline

Visual Input · Camera Stream

  1. Spatial Parsing

    Object · Position · Distance · Relation

  2. 3D Spatial Model

    Scene Geometry / Spatial Graph

  3. Persistent Spatial Memory

    Object Tracking / Temporal Memory

  4. Context & Relevance Engine

    User State × Environment × Priority

  5. Adaptive Description

    What to say · When to say · How to say

Update · The decision feeds back into the spatial model and the memory

Frame → Caption

Recognise the current frame and turn it into a sentence

VS.

Space → Model → Memory → Decision

Model the space, remember it, judge, and deliver only what is needed

Recognition coverage

It tells four kinds of situation apart.

Hundreds of first-person scenes we filmed ourselves, laid out as round tiles scattered like a map. Similar scenes cluster into four groups, with example frames drawn from each group placed at the four corners.
Hundreds of first-person scenes arranged by recognition result. They fall into four clusters.
  1. Spatial tracking

    Spatial Tracking

    Walking along corridors and passages. People and walkable ground are given identifiers and followed as the same objects even as the view changes.

  2. People and obstacles

    People & Obstacle Detection

    Crowded shops and markets. Standing people, passing people, display units, and furniture are each identified separately.

  3. Scene understanding

    Scene Understanding

    Interiors full of glass walls and columns. Passable areas and blocked areas are separated surface by surface.

  4. Navigation in complex space

    Navigation in Complex Space

    Alleyways packed with signs and stalls. We deliberately included the conditions that make recognition hardest.

Demonstration

This is it running.

A three-minute film of what we have actually built. For anyone who cannot watch it, the scenes are written out below.

SightSpace Technical Demonstration · August 2026 · Length 3 minutes
Field record · Length 2 min 39 sec. Walking through a real gallery wearing the AI glasses, with each judgement appearing on the panel beside it as it happens. This is an unedited record, so it has no sound.

What the film shows

Opening
What we build
An introduction to the structure: turning what the glasses see into spatial information, and proposing the next action on top of it.
0:40
First, decide whether to speak
It looks first at the width available to pass and at nearby objects, then decides whether to speak at all. Most of the time it does not. When it does, it first composes a sentence of plain fact, then rewrites it in the tone and level of detail that suits the person.
1:00
It does not lose track of people
Moving things such as people and bags are given identifiers and followed as the same object even as the view changes. Filmed at a conference venue in Tokyo.
1:40
We tried it with our eyes closed
Moving through a real gallery with our eyes closed, relying on the glasses alone.
2:00
Rebuilding the space from many videos
Several recordings are combined to reconstruct the floor plan of the venue. Of 753 camera positions, 752 were registered.

Evidence

A record of the four running for real.

Spatial axis · World State

What is in front of me right now?

People, walkable ground, doorways, and glass walls are each identified as distinct things. Every one carries a mark showing whether the judgement is confirmed or merely inferred, so that nothing uncertain is ever spoken aloud.

Four scenes with coloured regions painted over them. People, walkable ground, doorways, and glass walls each carry a different colour and an identifier, with a confirmed or inferred label beside them.
Recognition results from indoor corridors, exhibition halls, and streets

Speech decision · Silence Policy

Should it speak right now?

There is always something to say: a person ahead, a doorway, a banner. We filter these through four stages. Has the passable width narrowed? Is this the first time seeing it? Is it dangerous? Was it just mentioned? Most candidates are filtered out here and end in silence.

Over 45 seconds walking a corridor, 23 moments allowed speech. One was spoken.
At a crowded entrance, three out of 66. Below, the objects the glasses have accumulated and the path already travelled.

Person axis · Persona

Does this person need to hear it?

From the same scene, three different sentences go out. For someone blind from birth, only what is needed to move, kept short. For someone who lost their sight later, the scene as it would have looked. For an older user, slower and warmer. It is not a change of tone; the choice of what to say diverges first.

Three guidance texts branching out of a single scene

Spatial reconstruction · Multi-video

What kind of space is this?

Five recordings, taken separately with glasses and with a drone, are merged so that every camera position lands in one coordinate frame. Of 753 positions, 752 were registered. From those points we fitted seven walls and eleven barrier posts into a plan, and shortlisted six likely locations for exhibited works.

Camera positions finding their place one by one
An analysis board in four parts. Top left, the camera paths of five recordings drawn in colour. Top right, the reconstructed floor plan of the venue. Bottom left, a displacement graph. Bottom right, a table of figures.
The final result. Top right is the venue floor plan rebuilt from video alone.

Logs

Screens we kept while building.

Screen recordings we kept while building. Not a highlight reel — these are the runs as they came out that day.

These are logs from before human review. As the warning banner on the footage says, we do not use the numbers on these screens as performance figures. The reviewed numbers are in the sections above.

What it measures while walking, and what it lets pass

What the glasses saw, and the decision made at that moment, on one screen.

  • What it measures each step, and the path behind it · 20 sec

    The top is what the glasses saw. Walkable floor is painted green, each person gets a box, and depth appears at the top right. The middle is the decision: the passable width is wide enough, so say nothing. The bottom draws the path already walked, split into open space, corridor, and narrow corridor.

  • What the glasses piled up over 45 seconds · 19 sec

    Even while saying nothing, the glasses keep looking. Text read and kinds of objects encountered accumulate, and the path walked is drawn on the right. Eyes only see the moment they are raised; the glasses see the whole walk. That difference is what the recap gives back later.

  • Even at a crowded entrance, it stays quiet · 1 min 40 sec

    The passable width has narrowed by half. It still says nothing, because it spoke a moment ago. Having spoken recently outranks having narrowed. Having a reason to speak is not the same as having to speak.

Cutting things apart, and keeping their numbers

People and objects are separated one by one and keep the same number as they pass. No faces are used.

  • People and their bags each get a number · 2 min

    People, handbags and chairs are cut into separate pieces and numbered. The same number follows a person as they overlap and walk past. The top left records how many were detected this frame, how many are visible, how many are new, and how many are ambiguous.

  • Confirmed and estimated are kept apart · 1 min 14 sec

    The same scene is split into confirmed and estimated. What is not certain is neither counted nor spoken. Classes we have decided to stay silent about drop off the list entirely.

The same screen, in several places

On the left is what the glasses saw; on the right is the World State at that moment. What is there (objects), what kind of space it is, and what might be about to happen (behaviour candidates) are written separately. We ran the same screen in different places to see what breaks first when conditions change.

  • A carpeted corridor · 1 min

    A corridor with several people passing through. A busy floor pattern makes the walkable surface harder to hold on to. The bottom right also explains what each on-screen marking means.

  • An indoor shopping street with bright signs · 1 min

    Lit signage and a lot of glass. Text reads well here, but glass and reflections unsettle the walkable-floor judgement.

  • Dark surroundings, bright signs · 4 min 8 sec

    Everything is dark except the signage. We ran this one long to see how depth estimation wavers where the brightness range is extreme.

  • A booth covered in text · 4 min 11 sec

    A booth with a laptop and a display panel. Here a behaviour candidate — someone might linger in this spot — actually fires. It was the only one of the five places where that happened.

  • A crowded exhibition hall · 5 min

    The busiest condition of the set. Five minutes without a cut, to see whether the screen holds while people keep moving in and out.

Rebuilding a space from several videos

How separately shot footage is brought into one coordinate system.

  • From points to a floor plan · 24 sec

    Camera positions are aligned to build a point cloud, several videos are laid onto the same coordinates, and then walls, openings, artworks and routes are attached to make a plan. The result of this process, with its numbers, is in demonstration 04 above.

In the field

Where we filmed.

Scale

We hold agreements with three museums and two centres for blind people in Korea, and have filmed on site at more than fifteen indoor venues. Exhibition halls, conferences, shops, and streets were captured from a first-person view. We wrote our own specification covering thirty classes of spatial object, reviewed 2,037 frames sampled every second, then reviewed a further 6,822 frames at six-frame intervals.

Diversity

We filmed wearing five different models of AI glasses, and placed drone footage into the same coordinate frame. Bright exhibition halls, dark conference venues, crowded entrances, shops full of glass walls: we deliberately chose the conditions that make recognition harder.

Privacy

All processing finishes on the device. Recorded video is never uploaded to a server. We do not track faces or gaze, and we keep no per-person record of who looked at what, or for how long.

  • A conference hall with people and objects painted in colour and labelled with identifiers.
  • Products and signage being recognised inside a shop.
  • A first-person view walking into the dark entrance of an exhibition hall.
  • Viewing an exhibition in a gallery while wearing the AI glasses.
  • A first-person view walking along an indoor corridor.
  • A three-dimensional plan of the venue reconstructed from several videos.
  • Judgement results displayed in front of an artwork in an exhibition hall.
  • A sculpture being recognised with its state displayed.
  • The passable width of a corridor floor marked with a yellow line.

One memory, three uses

The same memory, read in three directions.

Experience Memory

Order travelled · Places lingered · Confirmed · Missed

  1. While walking

    Guide

    Says only what is needed, in real time

  2. After a trip

    Recap

    Retraces what one person confirmed and what they missed

  3. After many visitors

    Report

    Tells the venue operator where people get stuck

These are not three separate products. They are one memory, read in three directions.

Where it runs

On the device, with no internet.

It runs on the device, with no internet connection. It works in underground passages and inside buildings where the signal is weak, and the recorded video never leaves the device.

Jinseorang Madison Chung holding out a palm-sized computing board and Juyeon Lee holding out a pair of AI glasses toward the camera.
On the left, the computing board that runs the glasses without an internet connection

Measured figures

Measured on real footage and real hardware.

Measured on 300 seconds of first-person footage filmed indoors, and on the laptop and edge device we actually use.

14 of 601

Moments where speech was possible in 300 seconds, and how many were actually spoken

0.13 sec

Time to read one scene. About the length of a blink

1.9 GB

Size of the model that sits on the device. Roughly a few hundred photographs

150 combinations

Unseen combinations of user conditions where instructions were still followed