Collected sources and patterns will appear here. Add from search or the patterns library.
ParquetDataset -> Stream<Row>
Shuffle Parquet row groups first, then shuffle individual rows within a sliding-window memory buffer.
Problem it solves
Global shuffling of massive disk-backed datasets exceeds memory limits, while naive local shuffling causes highly correlated samples.
Consumes
Emits
The real projects this mechanism was found in. Attribution is the point — this is how the best teams actually do it.