Skip to content

Migrating to Chunklet 3.x.x

Whether you're coming from v2 or (bravely) from v1, this guide gets you onto the v3 API.

Automated migration checker

I wrote a script that scans your code for old patterns. It'll point out exactly what needs changing.

curl -O https://raw.githubusercontent.com/speedyk-005/chunklet-py/main/audit_migration.py
python audit_migration.py /path/to/your/project

Coming from v2

The v2-v3 jump is small: no API renames, just a couple of removals. If you're already on the unified chunk_text/chunk_texts/split_text API, you're unaffected by the deprecation removals.

Removed v2.2.0 aliases

The names deprecated in v2.2.0 are gone in v3. If you were still using them, here's the mapping:

  • SentenceSplitter.split() => split_text()
  • DocumentChunker.chunk() => chunk_text() / chunk_file()
  • DocumentChunker.batch_chunk() => chunk_texts() / chunk_files()
  • CodeChunker.chunk() => chunk_text() / chunk_file()
  • CodeChunker.batch_chunk() => chunk_texts() / chunk_files()
  • PlainTextChunker public import => DocumentChunker.chunk_text()

If you were already calling the chunk_text/chunk_texts/split_text methods, nothing changes for you.

DotDict serializers trimmed, to_dict renamed

DotDict carried to_json(), to_yaml(), to_toml(), and to_msgpack() for compatibility with the old python-box API. None were ever used inside chunklet and only to_dict() was called anywhere at all, so the other four are removed. Use the standard library instead — the names and the output are the same:

chunk.to_dict()              # -> {"content": "...", ...} 
chunk.to_json()              # -> '{"content": "...", ...}'
import json
json.dumps(chunk.to_std_dict())

to_dict() is now to_std_dict() on both DotDict and DotList. The old name was misleading: DotDict already is a dict, so to_dict() never meant "give me a dict".

You probably don't need it

DotDict subclasses dict, so json.dumps(chunk) works directly, and chunk == {"content": "..."} compares equal to a plain dict. You only need to_std_dict() when a consumer insists on exact standard types, e.g., yaml.dump() emits !!python/object tags for dict subclasses unless you convert first.

Custom sentence splitters are gone

Removed in v3.0.0

Custom splitters were a v2 feature: the old custom_splitters constructor parameter was replaced by a global custom_splitter_registry in v2, and both were removed entirely in v3.0.0.

SentenceSplitter now always uses its built-in language handlers, falling back to a universal rule-based splitter for unsupported languages. If you were relying on a custom splitter for a specific language, open a feature request or split that language manually.

Custom processor registry is now instance-based

In v2, custom processors lived on a global custom_processor_registry singleton shared across your whole app. In v3 that global is gone, you now create a CustomProcessorRegistry() yourself and pass it to a DocumentChunker via processor_registry. Each registry is independent, so registrations are scoped to the chunker you attach it to.

Fix:

from chunklet.document_chunker import DocumentChunker, custom_processor_registry


@custom_processor_registry.register(".json", name="MyJSONProcessor")
def my_json_processor(file_path: str) -> tuple[str, dict]: ...


chunker = DocumentChunker()
from chunklet.document_chunker import CustomProcessorRegistry, DocumentChunker

registry = CustomProcessorRegistry()


@registry.register(".json", name="MyJSONProcessor")
def my_json_processor(file_path: str) -> tuple[str, dict]: ...


chunker = DocumentChunker(lang="en", processor_registry=registry)

Scope your registries

Share a single CustomProcessorRegistry() instance across chunkers only when you actually want them to share the same custom processors.

lang="auto" is no longer the default

In v2, lang defaulted to "auto" and py3langid was a hard dependency; it was always installed. In v3, lang is required (no default), and py3langid is now an optional extra called [lang-detect].

If you were relying on automatic language detection, you need to:

  1. Pass lang="auto" explicitly (it's no longer implicit).
  2. Install the extra: pip install 'chunklet-py[lang-detect]'

If you only ever used specific language codes like lang="en", you don't need the extra; the default install is enough.

No auto-detection warning for you

In v2, SentenceSplitter warned on first use with lang="auto" ("Consider setting the lang parameter to a specific language"). That warning is removed in v3. Auto-detection still works, and the detected language and confidence are still logged at verbose level. If you relied on that warning, note that lang is now a required argument, so you already have it in hand: check lang == "auto" yourself and emit your own warning outside the library.

chunker = DocumentChunker()
chunks = chunker.chunk_text(text)  # lang defaulted to "auto"
chunker = DocumentChunker(lang="auto")
chunks = chunker.chunk_text(text)

And install the extra:

pip install 'chunklet-py[lang-detect]'

show_progress now defaults to False

In v2, batch methods (chunk_texts, chunk_files) showed a progress bar by default. In v3, show_progress defaults to False everywhere. Pass show_progress=True explicitly if you want the bar back.

chunks = list(chunker.chunk_files(paths))  # progress bar shown
chunks = list(chunker.chunk_files(paths, show_progress=True))

Language detection moved to common

Language detection is no longer a method on SentenceSplitter. It's now a standalone function at chunklet.common.lang_detection.detect_top_language(), so you can call it without constructing a splitter.

This affects two ways you may have used it before:

  • The legacy chunklet.utils.detect_text_language() from v1/v2.
  • SentenceSplitter.detected_top_language() from v2.
from chunklet.sentence_splitter import SentenceSplitter

splitter = SentenceSplitter(lang="auto")
lang_code, confidence = splitter.detected_top_language(text)
from chunklet.common.lang_detection import detect_top_language

lang_code, confidence = detect_top_language(text)

It still needs py3langid, so install the extra if you haven't:

pip install 'chunklet-py[lang-detect]'

offset is gone

The offset parameter is removed. It skipped the first N sentences before chunking, and returned nothing when the offset exceeded the sentence count. It's gone from DocumentChunker, PlainTextChunker, and the --offset CLI flag.

offset sliced the chunker's internal sentence list before grouping, and that list is not exposed by any public API, so there is no drop-in replacement.

The only exact equivalent is to split the text yourself and slice at a sentence boundary, using a splitter that reports accurate character offsets. yasbd (already a core dependency) exposes these through BoundaryDetector.detect(), which yields the cumulative end offset of each sentence. Slicing the original string at one of those offsets reproduces exactly what offset selected, and re-joining split segments is not required. Mind the index: offset was a count of sentences to skip, while detect() returns a 0-indexed list of end offsets, so the boundary is at offset - 1.

chunks = chunker.chunk_text(text, offset=5)  # start at the 6th sentence
from yasbd import BoundaryDetector

offset = 5  # number of sentences to skip, as before
offsets = list(BoundaryDetector(lang="en").detect(text))
start = offsets[offset - 1] if offset else 0
chunks = chunker.chunk_text(text[start:])  # start at the 6th sentence

This splits sentences twice

Slicing at a boundary reported by yasbd means the text is split once in your code and once again inside the chunker, so it costs more than offset did. If you were only dropping a fixed preamble such as a license header or table of contents, slicing the text at a boundary you choose is cheaper and gives the same result.

Config moved to the constructor

Sizing and tuning parameters (max_tokens, max_sentences, max_section_breaks, overlap_percent, lang) used to be passed per call to chunk_text(), chunk_file(), chunk_texts(), chunk_files(), split_text(), and split_file(). They now live on the chunker/splitter instance, set once at construction and mutable as plain attributes. The same applies to CodeChunker (max_tokens, max_lines, max_functions) and SentenceSplitter (lang).

chunker = DocumentChunker(token_counter=...)
chunks = chunker.chunk_text(
    text,
    lang="auto",
    max_sentences=3,
    max_tokens=500,
    max_section_breaks=2,
    overlap_percent=20,
)
chunker = DocumentChunker(
    lang="auto",
    max_sentences=3,
    max_tokens=500,
    max_section_breaks=2,
    overlap_percent=20,
    token_counter=...,
)
chunks = chunker.chunk_text(text)

Extras renamed

The structured-document extra is renamed to struct-doc, and the document alias for it is gone. The old visualization extra is gone too.

pip install 'chunklet-py[structured-document]'
pip install 'chunklet-py[document]'
pip install 'chunklet-py[visualization]'
pip install 'chunklet-py[struct-doc]'
pip install 'chunklet-py[viz]'

Coming from v1

The v1-v2 rename guide (the old Chunklet class, mode, use_cache, batch_chunk, etc.) lives in the v2.x version of these docs. The short version: Chunklet became DocumentChunker, and the unified chunk_text/chunk_texts/split_text API replaced the old methods.


That's it. Go forth and migrate.

See CLI docs for the full breakdown.