RootLens·Real human work, made robot-ready.

Real human work is
training data.

RootLens is a physical-AI project born at the University of Osaka. We monetize work data from Japan's commercial workplaces and return the benefits, in many forms, to the workplaces themselves.

§01

Robots learn by watching people. Nobody is collecting the footage.

Teaching a robot to work takes a huge amount of video of people actually working. Just as ChatGPT learned language from a vast amount of text, robots learn human motion from a vast amount of video.

But usable work footage is in critically short supply. Billions of people do this kind of work every day, all over the world, yet there is no infrastructure that collects it as data. A precious resource is being thrown away unused.

The data robots can learn from is produced in the course of everyday operations, on site.

§02

We sell the work data collected on site, and the revenue funds services that help the workplace.

Work data has buyers. Robotics companies are searching worldwide for training footage. But it is not realistic for a shop or a factory to sell its own data.

So RootLens takes it all on, from filming to the sale. The revenue then goes back as a service shaped to what each workplace struggles with. For stores short on hands: staff who record while they work, plus a filming fee. For companies weighing automation: a free workplace diagnostic. More ways to give back will follow.

All we ask of the workplace: keep working as usual.

Global physical AI on one side, Japanese industry on the other. RootLens is building the crossing point.

§03

Services that help the workplace, concretely

Part-timers that pay you back

For any workplace that hires part-timers: restaurants, retail, light warehouse work. Hire a RootLens staffer wearing a headgear camera as a regular part-timer; we sell the footage recorded on shift as training data and return a filming fee for the hours recorded.

01 / 03
Work as usual
Staff wearing a headgear camera work part-time shifts. Dishwashing, prep, stocking, packing: any line of work is fine, as long as the headgear and the filming aren't in the way.
02 / 03
Footage becomes training data
The close-up work footage staff record on shift becomes valuable training data that teaches robots how to move. The data is never used for anything outside AI and robotics R&D.
03 / 03
The workplace gets paid
We pay a filming fee for the hours recorded: the workplace keeps the extra hands while holding labor costs down. The fee also goes to the staff themselves, adding a few hundred yen to their effective hourly wage.
THE MATH, PER HOUR
Wage the employer pays
¥1,200
Filming fee back
¥500
=
Effective cost
≈¥700/h

Example at a ¥1,200 hourly wage. The fee is paid per recorded hour. ¥500 is the current base rate; we raise it as data sales grow, aiming for zero effective cost.

Start with a workplace diagnostic

For companies considering robots or physical AI. We bring the capture equipment and app, record the work at your site, and deliver a diagnostic report, free of charge: which steps suit automation and what adoption would take.

It can be free because only the data you approve is sold as training data, and that revenue covers the diagnostic. All we ask of the site is the recording.

§04

Capture rigs and data specs

Robot training can't use just any video. It needs footage focused on the work itself, shot in real environments, from as many different sites as possible. RootLens data is collected in shops and workplaces that are actually in business. All filming takes place under a written agreement with each site, and a consent record is kept for every clip. Every clip is reviewed before delivery.

We grow the lineup rig by rig; one set is in operation today. See the sample data for format details.

SET 01 · IN OPERATION
iPhone Pro + headgear

An iPhone Pro mounted on the worker's head records first-person video while ARKit runs alongside. Requires iPhone 15 Pro or later.

Camera
Wide-angle camera, 1920×1440, 30 fps
Depth
LiDAR, 256×192, every frame, 16-bit (mm)
Confidence
3 levels (low / medium / high), every frame
Camera pose
Continuous 6DoF, ARKit VIO, consistent world frame
IMU
Accelerometer + gyroscope, nominal 100 Hz
Timestamps
Single nanosecond timebase across all sensors
Camera intrinsics
Included with every clip (RGB and depth)
Scene
ARKit mesh anchors, point cloud
Privacy processing
Face blurring (EgoBlur), no audio
SET 02 · IN PREPARATION
Smart glasses

A lightweight glasses-based rig is in preparation, with more sets to follow.

§05
Get in touch

Let's talk.

Staffing, automation consults, training data, or an offer to collaborate: we're always glad to hear from you.