Testing the Impact of Varying Chunk Sizes on Retrieval Performance
AI21 Labs conducted experiments using datasets like QMSum, NarrativeQA, and Seinfeld to test the effects of different chunk sizes. Each dataset was duplicated six times, with each duplication assigned a fixed chunk size ranging from 200 to 1,000 tokens. This methodology allowed for a systematic analysis of how the nature of a query dictates the optimal chunk size for retrieval.


