Skip to content

chunklet.self_tuning_chunker.self_tuning_chunker

A self-tuning chunker inspired by the "Adaptive Chunking" technique (Machine Learning Mastery, "Essential Chunking Techniques for Building Better LLM Applications"): chunking parameters are adjusted dynamically based on measured content characteristics instead of using fixed limits.

Classes:

  • SelfTuningChunker –

    Self-tune chunk boundaries for mixed text/code corpora from learned profiles.

SelfTuningChunker

SelfTuningChunker(
    lang: str,
    token_counter: Callable[[str], int] | None = None,
    hard_token_limit: int = 1024,
    initial_state: dict | None = None,
    verbose: bool = False,
)

Self-tune chunk boundaries for mixed text/code corpora from learned profiles.

The chunker maintains an Kaufman Adaptive Moving Average of structural metrics per content profile and uses the resulting estimates to size the chunk boundaries for each source it processes.

Key Features
  • Profile-based dispatch: classifies each source as code or document via heuristic.
  • Learning memory (KAMA profiles): persists per-profile stats via Kaufman Adaptive Moving Average.
  • Self-tuning limits: derives constraints from the learned profile each time instead of using fixed values.
  • Enriched chunk metadata: adds inferred_type to each chunk.

Initializes the SelfTuningChunker.

Parameters:

  • lang

    (str) –

    Language code (e.g., 'en', 'fr', 'auto'). Required; pass 'auto' to auto-detect per source (needs the [self-tuning] extra).

  • token_counter

    (Callable[[str], int] | None, default: None ) –

    Function that counts tokens in text. If None, token-based limits and max_tokens learning are disabled.

  • hard_token_limit

    (int, default: 1024 ) –

    Ceiling for the dynamically grown max_tokens.

  • initial_state

    (dict | None, default: None ) –

    Optional pre-calculated running average to seed the learned profiles (e.g. exported from a previous learned_state). Missing profiles or metrics fall back to the built-in defaults.

  • verbose

    (bool, default: False ) –

    Enable verbose logging.

Methods:

  • chunk_file –

    Chunks a single document/code from a given path using the constraints

  • chunk_files –

    Chunks multiple documents/codes from a list of file paths using the constraints

  • chunk_text –

    Chunks raw text content using the constraints configured at initialization.

  • chunk_texts –

    Chunks multiple text contents using the constraints configured

Attributes:

  • lang (str) –

    Get the chunking language code.

  • verbose (bool) –

    Get the verbosity status.

Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
@validate_input
def __init__(
    self,
    lang: str,
    token_counter: Callable[[str], int] | None = None,
    hard_token_limit: int = 1024,
    initial_state: dict | None = None,
    verbose: bool = False,
):
    """
    Initializes the SelfTuningChunker.

    Args:
        lang: Language code (e.g., 'en', 'fr', 'auto'). Required; pass 'auto'
            to auto-detect per source (needs the `[self-tuning]` extra).
        token_counter: Function that counts tokens in text.
            If None, token-based limits and max_tokens learning are disabled.
        hard_token_limit: Ceiling for the dynamically grown ``max_tokens``.
        initial_state: Optional pre-calculated running average to seed the learned
            profiles (e.g. exported from a previous ``learned_state``). Missing
            profiles or metrics fall back to the built-in defaults.
        verbose: Enable verbose logging.
    """
    self._verbose = verbose
    self._lang = lang
    self.token_counter = token_counter
    self.hard_token_limit = hard_token_limit

    self.histories = defaultdict(list)
    self.learned_state = self._merge_initial_state(initial_state)

    self._sentence_splitter = UniversalSplitter()

    # Initialize chunkers with sensible defaults; constraint attributes are
    # mutated per request from the learned params.
    self.document_chunker = DocumentChunker(
        lang=self._lang,
        max_sentences=7,
        token_counter=self.token_counter,
        verbose=self._verbose,
    )
    self.code_chunker = CodeChunker(
        max_lines=15,
        token_counter=self.token_counter,
        verbose=self._verbose,
    )

lang property writable

lang: str

Get the chunking language code.

verbose property writable

verbose: bool

Get the verbosity status.

chunk_file

chunk_file(
    path: str | Path,
    *,
    file_type: Literal["document", "code"] | None = None,
    _already_fitted: bool = False,
) -> list[DotDict]

Chunks a single document/code from a given path using the constraints configured at initialization.

Parameters:

  • path

    (str | Path) –

    The path to the document file.

  • file_type

    (Literal['document', 'code'] | None, default: None ) –

    Optional file type.

  • _already_fitted

    (bool, default: False ) –

    internal param to avoid double fitting.

Returns:

  • list[DotDict] –

    A list of DotDict objects, each representing a chunk with its content and metadata.

Raises:

Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
@validate_input
def chunk_file(
    self,
    path: str | Path,
    *,
    file_type: Literal["document", "code"] | None = None,
    _already_fitted: bool = False,
) -> list[DotDict]:
    """
    Chunks a single document/code from a given path using the constraints
    configured at initialization.

    Args:
        path: The path to the document file.
        file_type: Optional file type.
        _already_fitted: internal param to avoid double fitting.

    Returns:
        A list of `DotDict` objects, each representing a chunk with its content and metadata.

    Raises:
        InvalidInputError: If the input arguments aren't valid.
        FileNotFoundError: If provided file path not found.
        UnsupportedFileTypeError: If the file extension is not supported or is missing.
    """
    path = Path(path)
    file_type = file_type or self._detect_file_type(path)

    text_or_gen, metadata = self.document_chunker.extract_text_and_metadata(
        path, path.suffix
    )

    if isinstance(text_or_gen, str):
        return self.chunk_text(
            text_or_gen,
            file_type=file_type,
            base_metadata=metadata,
            _already_fitted=_already_fitted,
        )

    def fit_and_stream_spans(texts: Iterator) -> Iterator:
        for text in texts:
            self._fit(text, file_type)
            yield text, metadata, file_type

    chunk_func = partial(self.chunk_text, _already_fitted=True)

    return list(
        run_in_batch(
            func=chunk_func,
            iterable_of_args=fit_and_stream_spans(text_or_gen),
            iterable_name="texts",
            n_jobs=os.cpu_count() - 1,
            verbose=self.verbose,
        )
    )

chunk_files

chunk_files(
    paths: IterableOfPath,
    *,
    token_counter: Callable[[str], int] | None = None,
    separator: Any = None,
    n_jobs: Annotated[int, Field(ge=1)] | None = None,
    show_progress: bool = False,
    on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[DotDict, None, None]

Chunks multiple documents/codes from a list of file paths using the constraints configured at initialization.

This method is a memory-efficient generator that yields chunks as they are processed, without loading all documents into memory at once. It handles various file types.

Parameters:

  • paths

    (IterableOfPath) –

    A non-string iterable of paths to the document files.

  • token_counter

    (Callable[[str], int] | None, default: None ) –

    Optional token counting function.

  • separator

    (Any, default: None ) –

    A value to be yielded after the chunks of each source are processed. Note: None cannot be used as a separator.

  • n_jobs

    (Annotated[int, Field(ge=1)] | None, default: None ) –

    Number of parallel workers to use. If None, uses all available CPUs. Must be >= 1 if specified.

  • show_progress

    (bool, default: False ) –

    Display progress bar during processing. Defaults to False.

  • on_errors

    (Literal['raise', 'skip', 'break'], default: 'raise' ) –

    How to handle errors during processing. Can be 'raise', 'ignore', or 'break'.

Yields:

  • DotDict –

    DotDict object, representing a chunk with its content and metadata.

Raises:

Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
@validate_input
def chunk_files(
    self,
    paths: IterableOfPath,
    *,
    token_counter: Callable[[str], int] | None = None,
    separator: Any = None,
    n_jobs: Annotated[int, Field(ge=1)] | None = None,
    show_progress: bool = False,
    on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[DotDict, None, None]:
    """
    Chunks multiple documents/codes from a list of file paths using the constraints
    configured at initialization.

    This method is a memory-efficient generator that yields chunks as they
    are processed, without loading all documents into memory at once. It
    handles various file types.

    Args:
        paths: A non-string iterable of paths to the document files.
        token_counter: Optional token counting function.
        separator: A value to be yielded after the chunks of each source are processed.
            Note: None cannot be used as a separator.

        n_jobs: Number of parallel workers to use. If None, uses all available CPUs.
               Must be >= 1 if specified.
        show_progress: Display progress bar during processing. Defaults to False.
        on_errors: How to handle errors during processing. Can be 'raise', 'ignore', or 'break'.

    yields:
        `DotDict` object, representing a chunk with its content and metadata.

    Raises:
        InvalidInputError: If the input arguments aren't valid.
        FileNotFoundError: If provided file path not found.
        UnsupportedFileTypeError: If the file extension is not supported or is missing.
        MissingTokenCounterError: If `max_tokens` is provided but no `token_counter` is provided.
        CallbackError: If a callback function (e.g., custom processors callbacks) fails during execution.
    """

    def fit_and_stream_tuples(paths: IterableOfPath) -> Iterator:
        for path in paths:
            path = Path(path)
            file_type = self._detect_file_type(path)
            text_or_gen, metadata = self.document_chunker.extract_text_and_metadata(
                path, path.suffix
            )

            if isinstance(text_or_gen, str):
                text_or_gen = [text_or_gen]

            for text in text_or_gen:
                self._fit(text, file_type)
                yield text, metadata, file_type

    chunk_func = partial(self.chunk_text, _already_fitted=True)
    chunks = run_in_batch(
        func=chunk_func,
        iterable_of_args=fit_and_stream_tuples(paths),
        iterable_name="paths",
        separator=separator,
        n_jobs=n_jobs,
        show_progress=show_progress,
        on_errors=on_errors,
        verbose=self.verbose,
    )

    # HACK: Since a sentinel is always at the end of the gen,
    # and we are using itertools.chain, the last item of the chunks
    # might will be an empty one. e.g, [1, 2, 3] => 1-2, 2-3
    # The only work-around is to add a mock to the end.
    chunks = chain(chunks, [DotDict({"metadata": {"source": ""}})])

    prev = None
    for curr, next in pairwise(chunks):
        if (
            curr == separator
            and prev is not None
            and prev.metadata.source == next.metadata.source
        ):
            continue

        if curr != separator:
            prev = curr

        yield curr

chunk_text

chunk_text(
    text: str,
    base_metadata: dict[str, Any] | None = None,
    file_type: Literal["document", "code"] | None = None,
    _already_fitted: bool = False,
) -> list[DotDict]

Chunks raw text content using the constraints configured at initialization.

Parameters:

  • text

    (str) –

    The raw text to fit and chunk.

  • base_metadata

    (dict[str, Any] | None, default: None ) –

    Optional dictionary to be included with each chunk.

  • file_type

    (Literal['document', 'code'] | None, default: None ) –

    Optional file type.

  • _already_fitted

    (bool, default: False ) –

    internal param to avoid double fitting.

Returns:

  • list[DotDict] –

    A list of DotDict objects, each representing a chunk.

Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
@validate_input
def chunk_text(
    self,
    text: str,
    base_metadata: dict[str, Any] | None = None,
    file_type: Literal["document", "code"] | None = None,
    _already_fitted: bool = False,
) -> list[DotDict]:
    """Chunks raw text content using the constraints configured at initialization.

    Args:
        text: The raw text to fit and chunk.
        base_metadata: Optional dictionary to be included with each chunk.
        file_type: Optional file type.
        _already_fitted: internal param to avoid double fitting.

    Returns:
        A list of `DotDict` objects, each representing a chunk.
    """

    file_type = file_type or self._detect_file_type(text)
    if not _already_fitted:
        self._fit(text, file_type)

    learned_state = self.learned_state[file_type]
    if file_type == "code":
        self.code_chunker.token_counter = self.token_counter
        self.code_chunker.max_tokens = (
            min(self.hard_token_limit, learned_state["max_tokens"])
            if self.token_counter
            else None
        )
        self.code_chunker.max_lines = learned_state["max_lines"]
        self.code_chunker.max_functions = round(learned_state["max_functions"])

        chunks = self.code_chunker.chunk_text(
            text,
            token_counter=self.token_counter,
            strict=False,
            base_metadata=base_metadata,
        )
    else:
        self.document_chunker.max_tokens = (
            min(self.hard_token_limit, learned_state["max_tokens"])
            if self.token_counter
            else None
        )
        self.document_chunker.max_sentences = round(learned_state["max_sentences"])
        self.document_chunker.max_section_breaks = round(
            learned_state["max_section_breaks"]
        )

        chunks = self.document_chunker.chunk_text(
            text, token_counter=self.token_counter, base_metadata=base_metadata
        )

    for chunk in chunks:
        chunk["metadata"]["inferred_type"] = file_type

    return chunks

chunk_texts

chunk_texts(
    texts: IterableOfStr,
    *,
    base_metadata: dict[str, Any] | None = None,
    separator: Any = None,
    n_jobs: Annotated[int, Field(ge=1)] | None = None,
    show_progress: bool = False,
    on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[DotDict, None, None]

Chunks multiple text contents using the constraints configured at initialization.

Parameters:

  • texts

    (IterableOfStr) –

    A non-string iterable of texts to chunk.

  • base_metadata

    (dict[str, Any] | None, default: None ) –

    Optional dictionary to be included with each chunk.

  • separator

    (Any, default: None ) –

    A value to be yielded after the chunks of each source are processed.

  • n_jobs

    (Annotated[int, Field(ge=1)] | None, default: None ) –

    Number of parallel workers.

  • show_progress

    (bool, default: False ) –

    Display progress bar during processing. Defaults to False.

  • on_errors

    (Literal['raise', 'skip', 'break'], default: 'raise' ) –

    How to handle errors.

Yields:

  • DotDict –

    DotDict object, representing a chunk with its content and metadata.

Raises:

Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
@validate_input
def chunk_texts(
    self,
    texts: IterableOfStr,
    *,
    base_metadata: dict[str, Any] | None = None,
    separator: Any = None,
    n_jobs: Annotated[int, Field(ge=1)] | None = None,
    show_progress: bool = False,
    on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[DotDict, None, None]:
    """
    Chunks multiple text contents using the constraints configured
    at initialization.

    Args:
        texts: A non-string iterable of texts to chunk.
        base_metadata: Optional dictionary to be included with each chunk.
        separator: A value to be yielded after the chunks of each source are processed.
        n_jobs: Number of parallel workers.
        show_progress: Display progress bar during processing. Defaults to False.
        on_errors: How to handle errors.

    yields:
        `DotDict` object, representing a chunk with its content and metadata.

    Raises:
        InvalidInputError: If the input arguments aren't valid.
        UnsupportedFileTypeError: If the file extension is not supported or is missing.
    """

    def fit_and_stream_pairs(texts: IterableOfStr) -> Iterator:
        for text in texts:
            file_type = self._detect_file_type(text)
            self._fit(text, file_type)
            yield text, base_metadata, file_type

    chunk_func = partial(self.chunk_text, _already_fitted=True)

    yield from run_in_batch(
        func=chunk_func,
        iterable_of_args=fit_and_stream_pairs(texts),
        iterable_name="texts",
        separator=separator,
        n_jobs=n_jobs,
        show_progress=show_progress,
        on_errors=on_errors,
        verbose=self.verbose,
    )