Skip to content

chunklet.document_chunker._plain_text_chunker

Classes:

  • PlainTextChunker –

    A powerful text chunking utility offering flexible strategies for optimal text segmentation.

PlainTextChunker

PlainTextChunker(
    lang: str,
    max_tokens: int | None = None,
    max_sentences: int | None = None,
    max_section_breaks: int | None = None,
    overlap_percent: int = 20,
    token_counter: Callable[[str], int] | None = None,
    continuation_marker: str = "...",
    verbose: bool = False,
)

A powerful text chunking utility offering flexible strategies for optimal text segmentation.

Key Features
  • Flexible Constraint-Based Chunking: Segment text by specifying limits on sentence count, token count and section breaks or combination of them.
  • Clause-Level Overlap: Ensures semantic continuity between chunks by overlapping

at natural clause boundaries with Customizable continuation marker. - Multilingual Support: Leverages language-specific algorithms and detection for broad coverage. - Pluggable Token Counters: Integrate custom token counting functions (e.g., for specific LLM tokenizers). - Parallel Processing: Efficiently handles batch chunking of multiple texts using multiprocessing. - Memory friendly batching: Yields chunks one at a time, reducing memory usage, especially for very large documents.

Initialize The PlainTextChunker.

Constraints limit how each chunk is shaped; lang selects the language used for splitting. Limit constraints are validated at construction and on mutation. Args are documented on DocumentChunker.

Methods:

  • batch_chunk –

    Processes a batch of texts in parallel, splitting each into chunks.

  • chunk –

    Chunks a single text into smaller pieces based on constraints

Attributes:

  • verbose (bool) –

    Get the verbosity status.

Source code in src/chunklet/document_chunker/_plain_text_chunker.py
def __init__(
    self,
    lang: str,
    max_tokens: int | None = None,
    max_sentences: int | None = None,
    max_section_breaks: int | None = None,
    overlap_percent: int = 20,
    token_counter: Callable[[str], int] | None = None,
    continuation_marker: str = "...",
    verbose: bool = False,
):
    """
    Initialize The PlainTextChunker.

    Constraints limit how each chunk is shaped; `lang` selects the language
    used for splitting. Limit constraints are validated at construction and
    on mutation. Args are documented on `DocumentChunker`.
    """
    self._verbose = verbose
    self.token_counter = token_counter
    self.continuation_marker = continuation_marker
    self.max_tokens = max_tokens
    self.max_sentences = max_sentences
    self.max_section_breaks = max_section_breaks
    self.overlap_percent = overlap_percent
    self.lang = lang

    self._validate_constraints(
        max_tokens, max_sentences, max_section_breaks, token_counter
    )

    self.sentence_splitter = SentenceSplitter(lang=lang, verbose=self._verbose)
    self._initialized = True

verbose property writable

verbose: bool

Get the verbosity status.

batch_chunk

batch_chunk(
    texts: Iterable[str],
    *,
    token_counter: Callable[[str], int] | None = None,
    separator: Any = None,
    base_metadata: dict[str, Any] | None = None,
    n_jobs: int | None = None,
    show_progress: bool = False,
    on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[Any, None, None]

Processes a batch of texts in parallel, splitting each into chunks. Leverages multiprocessing for efficient batch chunking. Uses the constraints configured at initialization.

If a task fails, chunklet will now stop processing and return the results of the tasks that completed successfully, preventing wasted work.

Parameters:

  • texts

    (Iterable[str]) –

    A non-string iterable of input texts to be chunked.

  • token_counter

    (Callable[[str], int] | None, default: None ) –

    The token counting function. Required if max_tokens is set.

  • separator

    (Any, default: None ) –

    A value to be yielded after the chunks of each text are processed. Note: None cannot be used as a separator.

  • base_metadata

    (dict[str, Any] | None, default: None ) –

    Optional dictionary to be included with each chunk.

  • n_jobs

    (int | None, default: None ) –

    Number of parallel workers to use. If None, uses all available CPUs. Must be >= 1 if specified.

  • show_progress

    (bool, default: False ) –

    Display progress bar during processing. Defaults to False.

  • on_errors

    (Literal['raise', 'skip', 'break'], default: 'raise' ) –

    How to handle errors during processing. Defaults to 'raise'.

Yields:

  • Any –

    A DotDict object containing the chunk content and metadata, or any separator object.

Raises:

  • InvalidInputError –

    If texts is not an iterable of strings, or if n_jobs is less than 1.

  • MissingTokenCounterError –

    If max_tokens is provided but no token_counter is provided.

  • CallbackError –

    If an error occurs during sentence splitting or token counting within a chunking task.

Source code in src/chunklet/document_chunker/_plain_text_chunker.py
def batch_chunk(
    self,
    texts: Iterable[str],
    *,
    token_counter: Callable[[str], int] | None = None,
    separator: Any = None,
    base_metadata: dict[str, Any] | None = None,
    n_jobs: int | None = None,
    show_progress: bool = False,
    on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[Any, None, None]:
    """
    Processes a batch of texts in parallel, splitting each into chunks.
    Leverages multiprocessing for efficient batch chunking.
    Uses the constraints configured at initialization.

    If a task fails, `chunklet` will now stop processing and return the results
    of the tasks that completed successfully, preventing wasted work.

    Args:
        texts: A non-string iterable of input texts to be chunked.
        token_counter: The token counting function.
            Required if `max_tokens` is set.
        separator: A value to be yielded after the chunks of each text are processed.
            Note: None cannot be used as a separator.
        base_metadata: Optional dictionary to be included with each chunk.
        n_jobs: Number of parallel workers to use. If None, uses all available CPUs.
            Must be >= 1 if specified.
        show_progress: Display progress bar during processing. Defaults to False.
        on_errors: How to handle errors during processing.
            Defaults to 'raise'.

    Yields:
        A `DotDict` object containing the chunk content and metadata, or any separator object.

    Raises:
        InvalidInputError: If `texts` is not an iterable of strings, or if `n_jobs` is less than 1.
        MissingTokenCounterError: If `max_tokens` is provided but no `token_counter` is provided.
        CallbackError: If an error occurs during sentence splitting
            or token counting within a chunking task.
    """
    chunk_func = partial(
        self.chunk,
        base_metadata=base_metadata,
        token_counter=token_counter or self.token_counter,
    )

    yield from run_in_batch(
        func=chunk_func,
        iterable_of_args=texts,
        iterable_name="texts",
        n_jobs=n_jobs,
        show_progress=show_progress,
        on_errors=on_errors,
        separator=separator,
        verbose=self.verbose,
    )

chunk

chunk(
    text: str,
    *,
    token_counter: Callable[[str], int] | None = None,
    base_metadata: dict[str, Any] | None = None,
) -> list[DotDict]

Chunks a single text into smaller pieces based on constraints configured at initialization. Supports flexible constraint-based chunking, clause-level overlap, and custom token counters.

Parameters:

  • text

    (str) –

    The input text to chunk.

  • token_counter

    (Callable[[str], int] | None, default: None ) –

    Optional token counting function. Required for token-based modes only.

  • base_metadata

    (dict[str, Any] | None, default: None ) –

    Optional dictionary to be included with each chunk.

Returns:

  • list[DotDict] –

    A list of DotDict objects, each containing the chunk content and metadata.

Raises:

Source code in src/chunklet/document_chunker/_plain_text_chunker.py
def chunk(
    self,
    text: str,
    *,
    token_counter: Callable[[str], int] | None = None,
    base_metadata: dict[str, Any] | None = None,
) -> list[DotDict]:
    """
    Chunks a single text into smaller pieces based on constraints
    configured at initialization.
    Supports flexible constraint-based chunking, clause-level overlap,
    and custom token counters.

    Args:
        text: The input text to chunk.
        token_counter: Optional token counting function.
            Required for token-based modes only.
        base_metadata: Optional dictionary to be included with each chunk.

    Returns:
        A list of `DotDict` objects, each containing the chunk content and metadata.

    Raises:
        InvalidInputError: If any chunking configuration parameter is invalid.
        MissingTokenCounterError: If `max_tokens` is provided but no `token_counter` is provided.
        CallbackError: If an error occurs during sentence splitting or token counting within a chunking task.
    """
    max_tokens = self.max_tokens or sys.maxsize
    max_sentences = self.max_sentences or sys.maxsize
    max_section_breaks = self.max_section_breaks or sys.maxsize

    log_info(
        self.verbose,
        "Starting chunk processing for text starting with: {}.",
        f"{text[:100]}...",
    )

    if not text.strip():
        log_info(self.verbose, "Input text is empty. Returning empty list.")
        return []

    self.sentence_splitter.lang = self.lang
    try:
        sentences = self.sentence_splitter.split_text(text)
    except Exception as e:
        raise CallbackError(
            f"An error occurred during the sentence splitting process.\nDetails: {e}\n"
            "💡 Hint: This may be due to an issue with the underlying sentence splitting library."
        ) from e

    if not sentences:
        return []

    chunks = self._group_by_chunk(
        sentences,
        token_counter=token_counter or self.token_counter,
        max_tokens=max_tokens,
        max_sentences=max_sentences,
        max_section_breaks=max_section_breaks,
        overlap_percent=self.overlap_percent,
    )

    # Note: We use DeterministicSpanFinder because sentence splitter may modify text
    # (e.g., normalize whitespace, fix encoding), making exact span tracking difficult.
    span_finder = DeterministicSpanFinder(text)
    return self._create_chunks(chunks, base_metadata or {}, span_finder)