Computer vision begins with individual observations: a person appears at a position, moves through the image and may become partially hidden by another visitor or an architectural element. The visual system needs a continuous spatial signal. The work between those two states determines whether the response feels calm and immediate or unstable.
The useful question is therefore behavioural: which movement matters, how quickly should the installation answer, and how should it behave when the scene becomes ambiguous?
From detections to stable tracks
A detector describes what appears in the current frame. An interactive installation usually needs continuity across many frames: a persistent track, a position within the relevant spatial zone and a confidence level that can guide the response.
The tracking layer turns changing observations into a stable spatial signal. It maintains continuity through brief occlusion and reduces visible jitter while preserving the response time required by the content.
This balance follows the content behaviour. A slow field of light can accept more smoothing than a direct gesture. A visitor approaching a media surface may work through stable proximity bands. The spatial event defines the required signal.
Site conditions shape the sensing design
Camera position, lens, field of view and processing method are part of the spatial design. Mounting height changes occlusion. Viewing angle changes how movement maps into the image. Ambient light affects contrast and exposure. Crowd density affects how long a person remains individually readable.
The sensing study therefore begins with the site and the intended behaviour. It identifies the active zones, likely paths, expected density, available mounting positions and the response time the content requires. Tests under representative light and movement conditions reveal the usable operating range.
Defined behaviour under uncertainty
Tracking quality changes as a space fills, light shifts or people overlap. A dependable installation treats confidence as a design input. The visual response can soften, retain its previous state for a short period or return gradually to an autonomous behaviour.
Multiple sensing layers can provide additional evidence where the site and interaction justify the added complexity. The purpose of fusion is specific: one input confirms or complements another at a defined decision point. The system architecture should also retain a clear operating state when one input becomes unavailable.
Nespresso New York
At the Nespresso flagship in New York, Refractiv designed the interactive system and sensor layout for a generative media wall. The sensing layer combines computer vision with multiple LiDARs and translates visitor proximity and movement into forces within a fluid simulation.
The composition remains active throughout the day. Presence changes the local behaviour of the visual field, while the authored system retains a coherent coffee-and-cream material language. Refractiv's scope included system design, hardware specification, implementation and calibration, connecting the sensing model to the final spatial response.
Designing the signal before choosing the sensor
A people-tracking system becomes useful when its output matches the intended behaviour of the work. Defining the spatial event, response time, confidence model and fallback state creates the criteria for hardware selection, software development and site validation.
For a sensing feasibility study or the development of a spatial response system, contact Refractiv.
Related project: Nespresso New York interactive video wall. Related capability: Sensing and Spatial Response.