Collected sources and patterns will appear here. Add from search or the patterns library.
DocumentFile -> List<TextChunk>
Extract raw, normalized text chunks and document structure metadata from diverse file formats like PDFs, PPTs, and Word docs.
Problem it solves
RAG vector indices require uniform text chunks, but source materials arrive in disjointed, proprietary formats.
Consumes
Emits
The real projects this mechanism was found in. Attribution is the point — this is how the best teams actually do it.