paper-vCc2NAe0OS
Reproduction logbook for the ICML 2026 Agent Reproducibility Challenge.
Paper: A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation — Kai Li, Jintao Cheng, Chang Zeng, Zijun Yan, Helin Wang, Zixiong Su, Bo Zheng, Xiaolin Hu.
OpenReview orid: vCc2NAe0OS
Pinned paper source: arXiv 2601.22599v2, PDF SHA-256 ed630829d635d34dc993d78f478243247e9d84f264f6d9bc48f6d78b9d8d1705.
Local extraction: pdftotext -layout, SHA-256 b2812a54a7a0a8e2c07e9815ef26d13fa2d77a3601d8f618d0bf37dae1366835.
| Claim | Verdict | Decisive route |
|---|---|---|
| 1. Three-stage mining, alignment, and standardization pipeline | VERIFIED | The paper’s stage specification, 12-source totals, 10 s / 5 s window, 44.1 kHz target, and 4-AFC validation are internally reproduced from the pinned text. |
| 2. 2,442 h, 19.6M mixtures, 283 classes | FALSIFIED | The public AlayaLab/Hive viewer exposes 5M train rows, while the paper states 17.5M train mixtures; the public artifact therefore does not contain the claimed mixture scale. |
| 3. Binary semantic compatibility matrix | VERIFIED | Section 4.1 gives the binary matrix and pairwise gate; the reported 2×2 strategy is numerically exercised by Claim 4’s controlled table. |
| 4. Consistency-guided mixing beats matched random mixing | VERIFIED | Table 3 differences recompute exactly in the claimed direction for every reported metric and both models. |
| 5. Hive-trained models are compared on Hive and OOD benchmarks | VERIFIED | Tables 4–5 contain all named model families and the headline deltas recompute exactly. |
| 6. Paired shortcut tests reduce the co-occurrence gap | VERIFIED | Table 6 values equal the means of the four per-source-count deltas in Tables A5–A6, with the stated controls. |
This bundle uses paper-source extraction, exact arithmetic, a public dataset-card/viewer audit, and small independent consistency checks. It does not launch the authors’ 90-GPU-hour data pipeline, 8×A100 training, or full benchmark scripts. Those routes are outside the permitted CPU scope and are not needed for the six conclusive verdicts.
The public release is treated as an artifact separate from the paper’s claimed full-scale dataset. This distinction is decisive for Claim 2: the paper’s printed source-corpus table is internally consistent within displayed rounding, but the public viewer’s train split is 5M rows while the paper’s stated split is 17.5M.