Evidence workspace

See what the video supports—and what still needs asking

Practitioner speech populates the judgment record; visible observations and vision-model outputs remain in a separate visual-evidence layer. They can anchor precise questions without inventing meaning or correctness.

Episode, record and probe model available

15

Speech + on-screen text

08

Direct visual observations

539

Visual annotations + derived cues

20

Real model artifacts

07

Interview probes

0:00 / 8:04

Source video

【町工場】ヤスリがけのコツ Part3【Milling machine】

家内事業町工場の日々 · 8:04

0399d97

Evidence contract

Only speech populates the judgment recordVisual observations propose the next questions

Structured record

Practitioner judgment record

Speech-backed only
01

Context

L字になっている部分をどう仕上げるか。

Audio not yet verified
02

Signals

面を感じながら掛ける。

Audio not yet verified

二方向からまっすぐ掛けると、角のところがきれいに掛からない。

Audio not yet verified
03

Criteria

角のところまでヤスリが掛かっている。

Audio not yet verified

角に掛け残しがない。

Audio not yet verified
04

Action

45度で当てて掛ける。

Audio not yet verified

掛け残しがないように掛け、二つの面をつなげる。

Audio not yet verified

ヤスリになっていない側を隣の面に当てて掛ける。

Audio not yet verified
05

Reasoning

ヤスリになっていない側を当てると、隣の側を傷つけず、掛けたい側だけを削れるため。

Audio not yet verified
06

Exceptions

Not established in the source speech

02 · MULTIMODAL EVIDENCE

Visual evidence and machine observations

This separates literal video observations, authored attention, OpenCV-derived motion, and artifacts from models that actually ran. None of these alone establishes expert intent, rationale, quality, or safety.

Inspect overlays in Vision Lab
8Direct visual observations
539Visual annotations + derived cues
20Real model artifacts

Observations and model outputs support what is visible. What it means must be checked with a practitioner or separately cited existing knowledge.

01

Direct observations from the video

Literal descriptions of visible objects, orientation, possible contact, and movement only.

02

Timecoded visual annotations and derived cues

Attention points, motion trails, hand landmarks, and related candidates. Play any card at its source moment.

03

Artifacts from visual models that actually ran

semantic localizationgpt-5.6-sol

OpenAI · recorded-run

observations.json
hand poseMediaPipe Hand Landmarker

Google · mediapipe-0.10.35:fbc2a30080c3

hand-clip-01.json
segmentationSAM 2.1 Hiera Tiny

Meta · facebookresearch/sam2@2b90b9f5ceec907a1c18123530e92e794ad901a4

keyframes.jsoncurve-file-mask.pngcurve-support-glove-mask.pngcurve-workpiece-mask.pngl-angle-file-mask.pngl-angle-support-glove-mask.pngl-angle-workpiece-mask.pngstraight-file-mask.pngstraight-support-glove-mask.pngstraight-workpiece-mask.pngtool-show-glove-mask.pngtool-show-hand-tool-mask.pngtool-show-workpiece-mask.png
segmentationSAM 2.1 Hiera Tiny video predictor

Meta · facebookresearch/sam2@2b90b9f5ceec907a1c18123530e92e794ad901a4

262000-268000-4fps.json262000-272000-4fps.json376000-394000-4fps.json416000-426000-4fps.json84000-94000-4fps.json

Continue with this source

One source, five connected views

Switching views preserves the same source hash across instruction, analysis, skill structure, evidence, and human review.