chunklet.document_chunker._plain_text_chunker
Classes:
-
PlainTextChunker–A powerful text chunking utility offering flexible strategies for optimal text segmentation.
PlainTextChunker
PlainTextChunker(
lang: str,
max_tokens: int | None = None,
max_sentences: int | None = None,
max_section_breaks: int | None = None,
overlap_percent: int = 20,
token_counter: Callable[[str], int] | None = None,
continuation_marker: str = "...",
verbose: bool = False,
)
A powerful text chunking utility offering flexible strategies for optimal text segmentation.
Key Features
- Flexible Constraint-Based Chunking: Segment text by specifying limits on sentence count, token count and section breaks or combination of them.
- Clause-Level Overlap: Ensures semantic continuity between chunks by overlapping
at natural clause boundaries with Customizable continuation marker. - Multilingual Support: Leverages language-specific algorithms and detection for broad coverage. - Pluggable Token Counters: Integrate custom token counting functions (e.g., for specific LLM tokenizers). - Parallel Processing: Efficiently handles batch chunking of multiple texts using multiprocessing. - Memory friendly batching: Yields chunks one at a time, reducing memory usage, especially for very large documents.
Initialize The PlainTextChunker.
Constraints limit how each chunk is shaped; lang selects the language
used for splitting. Limit constraints are validated at construction and
on mutation. Args are documented on DocumentChunker.
Methods:
-
batch_chunk–Processes a batch of texts in parallel, splitting each into chunks.
-
chunk–Chunks a single text into smaller pieces based on constraints
Attributes:
-
verbose(bool) –Get the verbosity status.
Source code in src/chunklet/document_chunker/_plain_text_chunker.py
batch_chunk
batch_chunk(
texts: Iterable[str],
*,
token_counter: Callable[[str], int] | None = None,
separator: Any = None,
base_metadata: dict[str, Any] | None = None,
n_jobs: int | None = None,
show_progress: bool = False,
on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[Any, None, None]
Processes a batch of texts in parallel, splitting each into chunks. Leverages multiprocessing for efficient batch chunking. Uses the constraints configured at initialization.
If a task fails, chunklet will now stop processing and return the results
of the tasks that completed successfully, preventing wasted work.
Parameters:
-
(textsIterable[str]) –A non-string iterable of input texts to be chunked.
-
(token_counterCallable[[str], int] | None, default:None) –The token counting function. Required if
max_tokensis set. -
(separatorAny, default:None) –A value to be yielded after the chunks of each text are processed. Note: None cannot be used as a separator.
-
(base_metadatadict[str, Any] | None, default:None) –Optional dictionary to be included with each chunk.
-
(n_jobsint | None, default:None) –Number of parallel workers to use. If None, uses all available CPUs. Must be >= 1 if specified.
-
(show_progressbool, default:False) –Display progress bar during processing. Defaults to False.
-
(on_errorsLiteral['raise', 'skip', 'break'], default:'raise') –How to handle errors during processing. Defaults to 'raise'.
Yields:
-
Any–A
DotDictobject containing the chunk content and metadata, or any separator object.
Raises:
-
InvalidInputError–If
textsis not an iterable of strings, or ifn_jobsis less than 1. -
MissingTokenCounterError–If
max_tokensis provided but notoken_counteris provided. -
CallbackError–If an error occurs during sentence splitting or token counting within a chunking task.
Source code in src/chunklet/document_chunker/_plain_text_chunker.py
chunk
chunk(
text: str,
*,
token_counter: Callable[[str], int] | None = None,
base_metadata: dict[str, Any] | None = None,
) -> list[DotDict]
Chunks a single text into smaller pieces based on constraints configured at initialization. Supports flexible constraint-based chunking, clause-level overlap, and custom token counters.
Parameters:
-
(textstr) –The input text to chunk.
-
(token_counterCallable[[str], int] | None, default:None) –Optional token counting function. Required for token-based modes only.
-
(base_metadatadict[str, Any] | None, default:None) –Optional dictionary to be included with each chunk.
Returns:
-
list[DotDict]–A list of
DotDictobjects, each containing the chunk content and metadata.
Raises:
-
InvalidInputError–If any chunking configuration parameter is invalid.
-
MissingTokenCounterError–If
max_tokensis provided but notoken_counteris provided. -
CallbackError–If an error occurs during sentence splitting or token counting within a chunking task.