🤖 AI 资讯

· ·
← 返回列表

[D] How do you get preprocessed dataset of a paper [D]

Reddit r/MachineLearning2026-09-15 08:50:58扩散模型,招聘HR原文 ↗

Hi all,

I'm trying to reproduce a paper where the reported dataset statistics in Table 1 don't match what I get from the public raw data, even after implementing the preprocessing exactly as described.

I've tried all reasonable interpretations of the filtering described in the paper and the closest I can get is still an order of magnitude off for one of the datasets. The paper says "data available on request" — I emailed the authors and followed up once, no reply so far.

For those who've been in this spot:

  • Do you just keep the larger-but-valid version you can reproduce and document the mismatch?
  • Is it worth sampling to match the reported size or does that just create a different irreproducible dataset?
  • When do you escalate to the journal vs just waiting?

How have you successfully gotten preprocessed files from authors? Any etiquette around follow-ups or journal contacts that actually worked?

Thanks for any advice.

submitted by /u/Individual-Safety906
[link] [comments]