chunklet.self_tuning_chunker.self_tuning_chunker
A self-tuning chunker inspired by the "Adaptive Chunking" technique (Machine Learning Mastery, "Essential Chunking Techniques for Building Better LLM Applications"): chunking parameters are adjusted dynamically based on measured content characteristics instead of using fixed limits.
Classes:
-
SelfTuningChunker–Self-tune chunk boundaries for mixed text/code corpora from learned profiles.
SelfTuningChunker
SelfTuningChunker(
lang: str,
token_counter: Callable[[str], int] | None = None,
hard_token_limit: int = 1024,
initial_state: dict | None = None,
verbose: bool = False,
)
Self-tune chunk boundaries for mixed text/code corpora from learned profiles.
The chunker maintains an Kaufman Adaptive Moving Average of structural metrics per content profile and uses the resulting estimates to size the chunk boundaries for each source it processes.
Key Features
- Profile-based dispatch: classifies each source as code or document via heuristic.
- Learning memory (KAMA profiles): persists per-profile stats via Kaufman Adaptive Moving Average.
- Self-tuning limits: derives constraints from the learned profile each time instead of using fixed values.
- Enriched chunk metadata: adds inferred_type to each chunk.
Initializes the SelfTuningChunker.
Parameters:
-
(langstr) –Language code (e.g., 'en', 'fr', 'auto'). Required; pass 'auto' to auto-detect per source (needs the
[self-tuning]extra). -
(token_counterCallable[[str], int] | None, default:None) –Function that counts tokens in text. If None, token-based limits and max_tokens learning are disabled.
-
(hard_token_limitint, default:1024) –Ceiling for the dynamically grown
max_tokens. -
(initial_statedict | None, default:None) –Optional pre-calculated running average to seed the learned profiles (e.g. exported from a previous
learned_state). Missing profiles or metrics fall back to the built-in defaults. -
(verbosebool, default:False) –Enable verbose logging.
Methods:
-
chunk_file–Chunks a single document/code from a given path using the constraints
-
chunk_files–Chunks multiple documents/codes from a list of file paths using the constraints
-
chunk_text–Chunks raw text content using the constraints configured at initialization.
-
chunk_texts–Chunks multiple text contents using the constraints configured
Attributes:
Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
chunk_file
chunk_file(
path: str | Path,
*,
file_type: Literal["document", "code"] | None = None,
_already_fitted: bool = False,
) -> list[DotDict]
Chunks a single document/code from a given path using the constraints configured at initialization.
Parameters:
-
(pathstr | Path) –The path to the document file.
-
(file_typeLiteral['document', 'code'] | None, default:None) –Optional file type.
-
(_already_fittedbool, default:False) –internal param to avoid double fitting.
Returns:
-
list[DotDict]–A list of
DotDictobjects, each representing a chunk with its content and metadata.
Raises:
-
InvalidInputError–If the input arguments aren't valid.
-
FileNotFoundError–If provided file path not found.
-
UnsupportedFileTypeError–If the file extension is not supported or is missing.
Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
chunk_files
chunk_files(
paths: IterableOfPath,
*,
token_counter: Callable[[str], int] | None = None,
separator: Any = None,
n_jobs: Annotated[int, Field(ge=1)] | None = None,
show_progress: bool = False,
on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[DotDict, None, None]
Chunks multiple documents/codes from a list of file paths using the constraints configured at initialization.
This method is a memory-efficient generator that yields chunks as they are processed, without loading all documents into memory at once. It handles various file types.
Parameters:
-
(pathsIterableOfPath) –A non-string iterable of paths to the document files.
-
(token_counterCallable[[str], int] | None, default:None) –Optional token counting function.
-
(separatorAny, default:None) –A value to be yielded after the chunks of each source are processed. Note: None cannot be used as a separator.
-
(n_jobsAnnotated[int, Field(ge=1)] | None, default:None) –Number of parallel workers to use. If None, uses all available CPUs. Must be >= 1 if specified.
-
(show_progressbool, default:False) –Display progress bar during processing. Defaults to False.
-
(on_errorsLiteral['raise', 'skip', 'break'], default:'raise') –How to handle errors during processing. Can be 'raise', 'ignore', or 'break'.
Yields:
-
DotDict–DotDictobject, representing a chunk with its content and metadata.
Raises:
-
InvalidInputError–If the input arguments aren't valid.
-
FileNotFoundError–If provided file path not found.
-
UnsupportedFileTypeError–If the file extension is not supported or is missing.
-
MissingTokenCounterError–If
max_tokensis provided but notoken_counteris provided. -
CallbackError–If a callback function (e.g., custom processors callbacks) fails during execution.
Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 | |
chunk_text
chunk_text(
text: str,
base_metadata: dict[str, Any] | None = None,
file_type: Literal["document", "code"] | None = None,
_already_fitted: bool = False,
) -> list[DotDict]
Chunks raw text content using the constraints configured at initialization.
Parameters:
-
(textstr) –The raw text to fit and chunk.
-
(base_metadatadict[str, Any] | None, default:None) –Optional dictionary to be included with each chunk.
-
(file_typeLiteral['document', 'code'] | None, default:None) –Optional file type.
-
(_already_fittedbool, default:False) –internal param to avoid double fitting.
Returns:
-
list[DotDict]–A list of
DotDictobjects, each representing a chunk.
Source code in src/chunklet/self_tuning_chunker/self_tuning_chunker.py
chunk_texts
chunk_texts(
texts: IterableOfStr,
*,
base_metadata: dict[str, Any] | None = None,
separator: Any = None,
n_jobs: Annotated[int, Field(ge=1)] | None = None,
show_progress: bool = False,
on_errors: Literal["raise", "skip", "break"] = "raise",
) -> Generator[DotDict, None, None]
Chunks multiple text contents using the constraints configured at initialization.
Parameters:
-
(textsIterableOfStr) –A non-string iterable of texts to chunk.
-
(base_metadatadict[str, Any] | None, default:None) –Optional dictionary to be included with each chunk.
-
(separatorAny, default:None) –A value to be yielded after the chunks of each source are processed.
-
(n_jobsAnnotated[int, Field(ge=1)] | None, default:None) –Number of parallel workers.
-
(show_progressbool, default:False) –Display progress bar during processing. Defaults to False.
-
(on_errorsLiteral['raise', 'skip', 'break'], default:'raise') –How to handle errors.
Yields:
-
DotDict–DotDictobject, representing a chunk with its content and metadata.
Raises:
-
InvalidInputError–If the input arguments aren't valid.
-
UnsupportedFileTypeError–If the file extension is not supported or is missing.