Overview

Welcome to the programmatic interface! This is where you integrate Chunklet-py's chunking capabilities directly into your Python apps.

  • Sentence Splitter


    Splits text into sentences across 60+ languages with automatic language detection and complex structure handling.

    Great for preparing clean text data for NLP tasks, LLMs, or any application that needs accurate sentence boundaries.

    Learn More

  • Document Chunker


    Transforms plain text and diverse document formats (.pdf, .docx, .epub, .eml, .pptx, .txt, .tex, .html, .hml, .md, .rst, .rtf, .odt, .csv, and .xlsx) into sized chunks with composable constraints and overlap for LLM and embedding pipelines.

    Great for RAG systems, document analysis, or any workflow that needs chunk size control.

    Learn More

  • Code Chunker


    Chunks source code while preserving logical structure and maintaining code semantics across functions, classes, and modules.

    Language-agnostic and lightweight, great for code understanding, generation, analysis, documentation, and AI model training.

    Learn More

  • Self-Tuning Chunker


    Self-tuning chunks for mixed text and code corpora. Classifies each source as document or code, learns per-profile structural statistics via a Kaufman Adaptive Moving Average (KAMA), and sizes chunk boundaries to match your content without manual constraint tuning.

    Perfect for heterogeneous corpora, evolving codebases, and anyone tired of hand-picking limits.

    Learn More

  • Text Chunk Visualizer


    Interactive web interface for real-time chunk visualization, parameter tuning, and exploring chunking results with live feedback.

    Great for experimenting with chunking strategies and comparing different settings.

    Learn More

Pick a card below to get started! 📇