Under Review at the ACM Multimedia 2026 Dataset Track

IMPACT: A Dataset for Multi-Granularity Human Procedural Action Understanding
in Industrial Assembly

Di Wen1, Zeyun Zhong1, David Schneider1, Manuel Zaremski1, Linus Kunzmann1

Yitian Shi1, Ruiping Liu1, Yufan Chen1, Junwei Zheng1,2, Jiahang Li1

Jonas Hemmerich1, Qiyi Tong3, Patric Grauberger1, Arash Ajoudani3, Danda Pani Paudel4

Sven Matthiesen1, Barbara Deml1, Jürgen Beyerer1, Luc Van Gool4, Rainer Stiefelhagen1, Kunyu Peng1,4,*

1Karlsruhe Institute of Technology, Germany 2ETH Zurich, Switzerland 3Italian Institute of Technology, Italy 4INSAIT, Sofia University, Bulgaria

* Corresponding author: kunyu.peng@kit.edu

A synchronized five-view RGB-D dataset and benchmark for industrial assembly under viewpoint shift, partial observability, multi-route execution, and anomaly recovery.

Benchmark code, task protocols, and released baselines are maintained in the GitHub repository.

Assembly Task and Recording Setup

IMPACT is collected around real angle grinder assembly and disassembly with synchronized ego and exo RGB-D, gaze, audio, and cognitive metadata.

Two angle grinder configurations and associated assembly parts used in IMPACT.
Manual book showing standardized installation and dismantling order.
Data acquisition pipeline for the IMPACT dataset.
IMPACT data acquisition workstation.
Trials 112 trials
Participants 13 subjects
Duration 39.5 video hours
Views 1 ego + 4 exo RGB-D

Dataset Overview

IMPACT comprises 112 trials from 13 participants and 39.5 video hours of synchronized industrial assembly data.

Dataset statistics summary for the IMPACT benchmark.
Assembly object

Two physical product configurations

Two angle grinder configurations provide a natural cross-configuration generalization setting.

Execution structure

Partial-order procedural execution

Valid trajectories follow a prerequisite graph rather than a single rigid sequence.

Supervision

Anomaly and compliance supervision

Explicit anomaly recovery supervision is paired with compliance phases and a six-category anomaly taxonomy.

Human metadata

Trial-level cognitive workload

Each trial is paired with NASA-TLX workload measurement.

Benchmark Tasks

Temporal understanding task family: TAS-S, TAS-BL, and TAS-BR.
Cross-view understanding task family: CV-TA, CV-SMR, and CV-SMC.
Action forecasting task family: AF-S and AF-L.
State and reasoning task family: PSR, ASR, PPR, and ATR.

Download

Versioned archives provide annotations, multi-view media, ego audio and eye tracking, released features, and a compact quick-start sample.

Type Content
Benchmark repository Official protocols, splits, mappings, task entry points, and evaluation code.
Full data release Annotations, five-view RGB, four-view depth, ego audio and eye tracking, and released feature packs.
Quick-start sample Three complete executions with synchronized media, features, and task annotations in IMPACT-v1.1-sample.zip.
Integrity metadata Archive byte sizes and SHA-256 checksums are published with each v1.1 mirror.

Release Update

Versioned annotation updates are documented with reproducible release notes.

Released · Jul 2026

IMPACT v1.1

We further improved the released annotations and videos. All 560 TAS-B annotation files were revalidated, and 117 of 560 RGB videos (20.9%) were updated. The refined release is now available on Google Drive and Hugging Face.

Annotation Schema

Aligned supervision spans interaction dynamics, procedural structure, compliance, and state evolution.

Annotation demo for the IMPACT dataset.
Data annotation workflow for the IMPACT dataset.

Fine-grained actions

Hand-specific labels capture coordinated bimanual interaction.

Procedural structure

Coarse steps and completion events connect segmentation to graph-based reasoning.

Assembly states

Component-wise ternary states support direct ASR and indirect PSR.

Compliance and anomalies

Compliance phases are paired with anomaly categories for diagnosis.

Benchmark

Released methods are organized by task family, with detailed protocols maintained in the task-specific pages.

Task coverage

  • Temporal understanding: TAS-S and TAS-BL/BR.
  • Cross-view understanding: CV-TA, CV-SMR, and CV-SMC.
  • Forecasting and reasoning: AF-S, AF-L, PSR, ASR, PPR-L/R, and ATR-L/R.

Metrics

  • Dense prediction metrics for temporal and compliance tasks.
  • Retrieval and classification metrics for cross-view tasks.
  • Task-specific forecasting, state, and procedural reasoning metrics.

Evaluation protocol

All splits are trial-level and co-assign synchronized views to prevent cross-view leakage. Full task definitions and reporting conventions are versioned with the benchmark code.

Resources

The repository is the operational layer for reproducible benchmarking and protocol inspection.

Citation and License

Stable citation metadata and release terms accompany the benchmark.

Manuscript & Citation

The manuscript is under review at the ACM Multimedia 2026 Dataset Track. Versioned author and dataset metadata are available in CITATION.cff; proceedings metadata will replace it after publication.

License

  • Repository-authored code: repository root LICENSE.
  • Dataset assets and documentation: LICENSE-DATA.
  • third_party/: per-directory provenance and license notices.

Acknowledgements

We gratefully acknowledge Lei Qi, Weitong Kong, Chen Zhang, Haiwen Sun, and Yuwei Hu for their valuable support throughout the data collection and annotation process for IMPACT.