Project Bumblebee

14,407 curated hours across four synchronized streams, measured on the axes that decide whether egocentric data is trainable.

  • Contactrobotics@deccan.ai
  • LicenseCommercial-training ready
  • FormatFHD–4K · 30–120 fps
  • DuplicatesZero, across all modalities
Representative egocentric capture from the household split.

Household egocentric

Curated hours
11,778
Clips
29,247
Participants
1,802
Tasks / domains
50 / 9

Factory egocentric

Curated hours
2,544
Clips
1,283
Participants
482
Tasks / families
52 / 10

Tri-camera

Curated hours
75
Clips (100×3)
300
Tasks
100
Viewpoints
Head + 2 wrists

Motion capture

Curated hours
9.7
Clips (72×4)
288
Tasks
18
Output
Metric 3D joints
Overview

What an hour of footage actually contains

Egocentric datasets are usually sold by total hours. For robotics that is the wrong headline. A manipulation policy needs hand-visible frames, a world model needs temporal structure, and an eval team needs certainty that no two clips are the same video twice. Bumblebee is built and reported against those properties.

14,407Curated hours across all four streams
31,118Released clips, each with versioned metadata
2,284Household and factory participants, plus 4 MoCap performers
100%Informed consent, commercial use included
The four streams

Each stream answers a different question

Two egocentric domains give breadth and precision. A tri-camera rig adds per-hand close-up observation. Marker-based motion capture supplies the metric skeleton everything else is validated against.

Household egocentric

11,778 hours across nine activity domains

Weighted toward the cleaning, organization, and cooking work that dominates domestic settings. The distribution is intentionally non-uniform: it follows where hands-on manipulation is densest rather than sampling domains evenly.

Household recorded hours by activity domain: cleaning 4,248; organization 2,387; cooking and food prep 1,820; laundry management 1,355; daily routines 612; assistance and navigation 538; preparation and occasions 381; hosting 258; home maintenance 179.
Activity domainClipsHours%
Cleaning12,0214,24836%
Organization6,6972,38720%
Cooking & food prep2,4191,82015%
Laundry management4,4881,35512%
Daily routines7096125%
Assistance & navigation1,9335385%
Preparation & occasions4453813%
Hosting3052582%
Home maintenance2301792%
Total29,24711,778100%

50 goal-directed tasks, grouped into these nine domains. Household footage also carries a tri-granular labeling of verb, object, and goal.

Factory egocentric

2,544 hours across ten operation families

Grouped by the manipulation skill each station demands rather than by location — all of this footage comes from one environment, the assembly floor. Three families carry 65% of factory hours: the tight-tolerance, high-repetition operations where policy performance matters most.

Factory recorded hours by operation family: component insertion and assembly 561; inspection and quality verification 549; precision soldering and joining 546; electrical test and debug 289; packaging and kitting 160; material processing and forming 155; machine tending and material handling 119; surface preparation and cleaning 64; marking and labeling 64; line supervision and unclassified 37.
Operation familyTasksHours%
Component insertion & assembly556022%
Inspection & quality verification954922%
Precision soldering & joining454621%
Electrical test & debug828811%
Packaging & kitting51606%
Material processing & forming61556%
Machine tending & handling51205%
Surface preparation & cleaning4653%
Marking & labeling3643%
Line supervision & unclassified3371%
Total522,544100%

52 distinct tasks recorded on active production lines. Per-task clip counts were not recorded for this split, so hours are reported without a clip column; the split comprises 1,283 clips in total.

Factory · deeper cut

The fifteen heaviest of 52 tasks

Top fifteen factory tasks by recorded hours, led by manual soldering and touch-up at 478 hours and manual insertion at 319 hours, down to marking on LCD at 33 hours.

Manual soldering and touch-up alone carries 478 hours — nearly a fifth of the factory split in a single operation. The concentration is what makes this data useful: deep coverage of the few operations that dominate a real line, rather than thin coverage spread evenly across all 52.

The remaining 37 tasks account for 579 hours. The full per-task listing, grouped by family, is in the report appendix.

Tri-camera

Head plus one camera per wrist

All three views are first-person. The wrist views resolve grasp and contact detail at a scale the head camera cannot, and they keep the hand in frame when the head view is blocked by the body or the object. This is the observation structure bimanual policies are usually conditioned on — hand-level resolution, not an external vantage point.

100Standardized tasks
3Synchronized views
300Clips (100 × 3)
75Curated hours

Each task is captured from all three views, so 100 tasks × 3 views = 300 clips. Hours are footage summed across the three views, at 15 minutes per view.

Motion capture

Metric ground truth, 18 task categories

Marker-based optical capture gives metric 3D joint trajectories — the reference that pixel-space pose estimation is validated against, and the basis for retargeting to humanoid and bimanual kinematic chains.

Motion-capture recorded minutes by task, led by packing at 106.3 minutes and arranging household items at 59.0 minutes, down to box transfer at 7.7 minutes.

18 tasks × 4 performers = 72 takes; each take is captured from 4 synchronized views, so 72 × 4 = 288 clips carrying 580 minutes. Team work and box transfer are two-performer interaction tasks.

What the data looks like

Frames from each capture configuration

One sample per configuration, at full width. Every frame below is from the released corpus.

Grid of representative egocentric frames spanning household scenes, actors, and lighting conditions.
Household egocentric Representative frames across household scenes, actors, and lighting conditions — the spread the lighting distribution describes numerically.
Grid of tri-camera frames. Each row is a bimanual manipulation task; the three frames in a row are the same time step from the head-mounted camera and one camera on each wrist.
Tri-camera · head + dual wrist Each row is a different bimanual task. The three frames within a row are the same time step, seen simultaneously from the head-mounted camera and one camera on each wrist.
Grid of motion-capture output. Rows show the three released streams: 3D pose render, solved skeleton, and RGB reference video, five samples each.
Motion capture · three released streams One row per released stream: 3D pose render (top), solved skeleton (middle), and RGB reference video (bottom), with five samples drawn from each.
Quality signals

Measured, not asserted

Hand visibility, lighting, and temporal integrity are measured on a random subset of 7,800 household clips drawn across all nine activity domains.

Hands

Hand visibility, quantified

Hand-visible frames (≥ 1 hand) 68%
Videos · both hands ≥ 20% of frames 45%
Videos · both hands ≥ 50% of frames 26%

Hand visibility is the most common filter teams apply to egocentric footage. Both-hands coverage, the signal that matters for bimanual work, is released per video at every threshold so a team can pick its own operating point and trade coverage against volume.

Lighting

Covering the deployment envelope

Scene-brightness distribution: share of videos by average luma normalized 0 to 1, a broad bell-shaped curve centred near mid-range with mean and median marked.

Scene brightness is measured directly as mean frame luminance rather than binned into coarse categories. The distribution is bell-shaped and centred near the middle of the normalized 0–1 range, corpus mean 0.44, with meaningful mass in both the darker and brighter tails. Deployed robots do not control their lighting, so spread is the property that matters — not a corpus shot under convenient studio light.

Temporal structure

Long-horizon material at scale

24Mean clip minutes
3,593Recordings ≥ 40 min
4,225Long-horizon hours
Histogram of clip durations across the corpus, with the mean and median marked. Duration histogram across all long-horizon recordings of at least 40 minutes, extending into multi-hour material.

The distribution is bimodal: short manipulation vignettes (88% of clips, mean 18 min) and long-horizon recordings (12% of clips, mean 71 min). The latter are an eighth of the clips but carry over a third of all hours.

Capture format

Native fidelity preserved

Share of parsed clips by native resolution: 1080p 70.2%, 720p 11.5%, 1080p portrait 7.4%, 4K UHD 5.6%, 720p portrait 1.5%, other 3.8%. Share of parsed clips by frame rate: 30 fps dominant, with substantial 60 fps and 50 fps and a 120 fps tail.

Clips ship at native capture resolution and frame rate rather than transcoded to a common target, so teams select the fidelity their pipeline expects and can study robustness to capture conditions instead of having it flattened away. The 120 fps tail resolves fast hand motion without blur.

Temporal integrity

Cuts are documented, not hidden

43%Clips free of jump cuts
0Undocumented cuts

Jump cuts and unannounced scene breaks corrupt sequence models. Clips that contain them ship with exact cut timestamps so downstream users can split or mask cleanly. Every cut identified in QC is documented with a timestamp rather than left implicit.

Uniqueness

Every clip is distinct

0Duplicates in the release
218Candidates removed pre-release

Duplicates inflate apparent scale, leak across eval splits, and let models overfit specific footage. Duplicate detection was run across the corpus and 218 candidates were removed before release.

  • Frame-exact synchronization across all four streams
  • Chain-of-custody documented from capture to release
  • No identifiable non-consenting individual in any frame
Task taxonomy

Two domains, two taxonomies

Household footage carries a tri-granular labeling — verb, object, goal. Factory tasks carry station-level identity instead, because a factory task name already denotes a fixed, repeatable operation on a known part.

510Verbs
2,140Objects
50Household goal tasks
9Household domains
52Factory tasks
10Factory operation families
Most-mentioned verb clusters
pick upwipeplacepressopen adjustfoldmovepourarrange closesmoothhandleinsertrotate removeinspectsprayturn onrub cutrepositionhangdipattach gathersortstirscoopspread
Most-mentioned object clusters
clothdrawerbowlclothesshirt bottlehandfloorcontainerstove pillowmopsinkbaglid cableboxbucketframeitem

Failure and recovery events — dropped objects, mis-grasps, corrections — are retained rather than filtered, across both splits. Curated imitation-learning data often removes them.

Curation

The QC behind the numbers

Every shipped clip cleared six stages in order, with attrition documented per stage so the material screened out is accounted for rather than hidden.

Funnel

Six stages, attrition documented

Six-stage curation funnel: ingestion validation 100%, automated quality screen 84%, VLM semantic verification 68%, human review 54%, duplicate elimination 44%, statistical audit and release 34%.

Per-stage attrition is documented and released to licensees. The ordering is deliberate: reported numbers change when the pipeline changes, not when a metric falls outside its target range.

Stages

What each stage rejects

  • 1Ingestion validation. Container, codec, timestamp monotonicity, per-frame decode.
  • 2Automated quality screen. Motion blur, sustained under- or over-exposure, incoherent ego-motion, audio corruption, encoding artifacts.
  • 3VLM semantic verification. A vision–language model checks scene-label agreement, hand-presence consistency, and safety-relevant properties. It is calibrated against human-in-the-loop review, and low-confidence cases escalate.
  • 4Human review. Uniform random-sample audit for calibration, plus targeted review of anything flagged upstream.
  • 5Duplicate elimination. Re-encodings, format conversions, and clips sharing substantial content despite differing framing or compression.
  • 6Statistical audit & release. Confidence intervals on every published metric.

Model choices, prompt structures, thresholds, and reviewer rubrics are not disclosed. Independent auditors are granted full access under NDA — the balance we think is right between verifiability and not handing over the methodology wholesale.

Delivery

What you actually receive

Every clip in all four streams ships with structured metadata, versioned alongside the data release. Schema changes stay backward-compatible within a major version.

Schema

Per-clip metadata

Clip {
  id          : uuid
  stream      : household | factory | tricam | mocap
  timestamps  : { capture_utc, duration_ms }
  scene       : { activity_domain | operation_family, task }
  lighting    : { avg_luma_norm }
  hands       : { pct_visible, both_hands_pct }
  task        : { verb, object, goal }
  quality     : { blur, exposure, ego_motion }
  integrity   : { jumpcut_count, cut_timestamps[], dedup_hashes }
  views       : { mount, sync_offset_ms }
  consent     : { participant_id, commercial_ok, jurisdiction }
}

Tri-camera metadata adds the mount position of each view and inter-view sync offsets. Motion-capture metadata adds per-take timecode, performer identity, approval status, and the synchronized reference-video pointer.

Coverage & licensing

Commercial-training ready

Core metadata 100%
Narration 100%
Hand-visibility stats 100%
Informed consent 100%
  • License supports commercial policy training and downstream model release
  • Informed consent documented for every participant, across all four streams
  • Bystander and PII handling per jurisdiction-appropriate policy
  • Filled datasheet and chain-of-custody with every release
  • Independent audit of consent, redaction, and provenance under NDA
What comes next

A living corpus

Collection continues beyond this release. The next deliverable goes one level deeper than clips and tasks, to the interactions themselves.

Immediate next release

Hand–object interaction analysis

A dedicated report over the released corpus: which objects are contacted, when contact begins and ends, how each grasp is formed and released, and how bimanual coordination is distributed across tasks. Contact structure is what separates footage a policy can learn manipulation from, and footage that merely shows a task being completed — and it is the axis teams most often have to annotate themselves before egocentric data becomes trainable.

Because the analysis runs over the corpus already released, it requires no re-collection and no change to the data licensees hold. It ships as an additive layer against the same versioned metadata schema.

01 — Richer modalities

More of what teams ask for

Instrumented tactile gloves for contact, grip force, and finger pose. UMI-style handheld grippers that record from the gripper's point of view rather than the wearer's, which transfers more directly to real arms. Eye-gaze tracking as a signal of attention and intent. Denser depth. Every new modality plugs into the same synchronization and metadata schema, so the corpus grows as one dataset rather than splitting into incompatible pieces.

02 — Long-tail environments

Coverage, not raw hours

Hours are no longer the bottleneck; rare situations are. We use the corpus's own metadata and clip embeddings to find environments that appear less often in our data than in the real world — cramped or dimly lit kitchens, shared and multi-generational households, festival preparation, aging appliances, unusual factory cells. Each gap becomes a collection brief, so rare settings are captured deliberately instead of by luck.

03 — The data pyramid

Human base, synthetic middle, robot top

The base is large-scale human data; Bumblebee sits at its high-quality end, first-person and synchronized and task-directed rather than scraped. The middle is simulation and generative augmentation, cheap to scale but only approximating real physics. The top is teleoperated and autonomous robot data, exact but expensive. Because every layer shares one task taxonomy and schema, data from one can train and validate models built on another.

04 — Adaptive distribution

Balance while collecting, not after

Today balance comes from collection briefs and after-the-fact curation. We are closing the loop: dashboards already show in near real time how much each task and scene contributes, and we are wiring those numbers back into task assignment. Once a category hits its target share, contributors are redirected to categories still short — so drift surfaces as an alert during collection instead of a surprise at release.

Access

We would be glad to hear from you

Bumblebee is available under a license that supports commercial policy training. A filled datasheet and chain-of-custody documentation are included with every release, and independent auditors are welcome to review consent, redaction, and pipeline detail under NDA.

License
Commercial-training ready
Cadence
Versioned · additive
Datasheet
Ships with release
Audit
Available under NDA