chunklet.sentence_splitter
Modules:
-
languages–This module contains the language sets for the supported sentence splitters.
-
sentence_splitter–
Classes:
-
SentenceSplitter–A robust and versatile utility dedicated to precisely segmenting text into individual sentences.
-
UniversalSplitter–Language-agnostic sentence boundary detector using regex patterns.
Functions:
-
detect_top_language–Detect the top language of the given text using py3langid.
-
log_info–Log an info message if verbose is enabled.
-
read_text_file–Read text file with automatic encoding detection.
-
validate_input–A decorator that validates function inputs and outputs
SentenceSplitter
A robust and versatile utility dedicated to precisely segmenting text into individual sentences.
Key Features: - Multilingual Support: Leverages language-specific algorithms and detection for broad coverage. - Fallback Mechanism: Employs a universal rule-based splitter for unsupported languages. - Intelligent Post-processing: Cleans up split sentences by filtering empty strings and rejoining stray punctuation.
Initializes the SentenceSplitter.
Parameters:
-
(langstr) –Language code (e.g., 'en', 'fr', 'auto').
-
(verbosebool, default:False) –If True, enables verbose logging for debugging and informational messages.
Methods:
-
split_file–Read and split a file into sentences.
-
split_text–Splits a given text into a list of sentences.
Source code in src/chunklet/sentence_splitter/sentence_splitter.py
split_file
Read and split a file into sentences.
Parameters:
-
(pathstr | Path) –Path to the file to read.
Returns:
-
list[str]–A list of sentences extracted from the file.
Source code in src/chunklet/sentence_splitter/sentence_splitter.py
split_text
Splits a given text into a list of sentences.
Parameters:
-
(textstr) –The input text to be split.
Returns:
-
list[str]–A list of sentences.
Examples:
>>> splitter = SentenceSplitter(lang="en")
>>> splitter.split_text("Hello world. How are you?")
['Hello world.', 'How are you?']
>>> splitter = SentenceSplitter(lang="fr")
>>> splitter.split_text("Bonjour le monde. Comment allez-vous?")
['Bonjour le monde.', 'Comment allez-vous?']
>>> splitter = SentenceSplitter(lang="auto")
>>> splitter.split_text("Hello world. How are you?")
['Hello world.', 'How are you?']
Source code in src/chunklet/sentence_splitter/sentence_splitter.py
UniversalSplitter
Language-agnostic sentence boundary detector using regex patterns.
A universal splitter using Unicode-aware regex patterns for any language.
Handles
- Unicode sentence terminators
- Numbered lists and headings
- Quoted sentences
- Line breaks and whitespace
Use cases
- Primary splitter for languages without dedicated support
- Fallback when language-specific splitters unavailable
Methods:
-
split–Splits text into sentences using rule-based regex patterns.
split
Splits text into sentences using rule-based regex patterns.
Parameters:
-
(textstr) –The input text to be segmented into sentences.
Returns:
-
list[str]–A list of sentences after segmentation.
Source code in src/chunklet/sentence_splitter/_universal_splitter.py
detect_top_language
Detect the top language of the given text using py3langid.
The LanguageIdentifier is built lazily on first use and cached for reuse.
Parameters:
-
(textstr) –The input text to detect the language for.
Returns:
-
str–A tuple of the ISO 639-1 language code and its confidence in
[0, 1]. -
float–Confidence depends on the
py3langidmodel, so treat it as -
tuple[str, float]–approximate rather than a threshold you can rely on across versions.
Raises:
-
ImportError–If py3langid is not installed.
Examples:
>>> lang, confidence = detect_top_language("This sentence is written in English.")
>>> lang, confidence > 0.8
('en', True)
>>> detect_top_language("Ceci est une phrase ecrite en francais.")[0]
'fr'
>>> code, confidence = detect_top_language("")
>>> round(confidence, 2)
0.01
Source code in src/chunklet/common/lang_detection.py
log_info
Log an info message if verbose is enabled.
This is a convenience function that only logs when verbose mode is enabled, avoiding unnecessary log output in production.
Parameters:
-
(verbosebool) –If True, logs the message; if False, does nothing.
-
–*argsPositional arguments passed to logger.info().
-
–**kwargsKeyword arguments passed to logger.info().
Example
log_info(True, "Processing file: {}", filepath) Processing file: /path/to/file log_info(False, "This will not be logged") (no output)
Source code in src/chunklet/common/logging_utils.py
read_text_file
Read text file with automatic encoding detection.
Parameters:
-
(pathstr | Path) –File path to read.
Returns:
-
str–File content.
Raises:
-
FileProcessingError–If file cannot be read.
Source code in src/chunklet/common/path_utils.py
validate_input
A decorator that validates function inputs and outputs
A wrapper around Pydantic's validate_call that catchesValidationError and re-raises it as a more user-friendly InvalidInputError.