Overview
Welcome to the programmatic interface! This is where you integrate Chunklet-py's chunking capabilities directly into your Python apps.
-
Sentence Splitter
Splits text into sentences across 60+ languages with automatic language detection and complex structure handling.
Great for preparing clean text data for NLP tasks, LLMs, or any application that needs accurate sentence boundaries.
-
Document Chunker
Transforms plain text and diverse document formats (
.pdf,.docx,.epub,.eml,.pptx,.txt,.tex,.html,.hml,.md,.rst,.rtf,.odt,.csv, and.xlsx) into sized chunks with composable constraints and overlap for LLM and embedding pipelines.Great for RAG systems, document analysis, or any workflow that needs chunk size control.
-
Code Chunker
Chunks source code while preserving logical structure and maintaining code semantics across functions, classes, and modules.
Language-agnostic and lightweight, great for code understanding, generation, analysis, documentation, and AI model training.
-
Self-Tuning Chunker
Self-tuning chunks for mixed text and code corpora. Classifies each source as document or code, learns per-profile structural statistics via a Kaufman Adaptive Moving Average (KAMA), and sizes chunk boundaries to match your content without manual constraint tuning.
Perfect for heterogeneous corpora, evolving codebases, and anyone tired of hand-picking limits.
-
Text Chunk Visualizer
Interactive web interface for real-time chunk visualization, parameter tuning, and exploring chunking results with live feedback.
Great for experimenting with chunking strategies and comparing different settings.
Pick a card below to get started! 📇