Skip to content

chunklet.sentence_splitter

Modules:

Classes:

  • SentenceSplitter –

    A robust and versatile utility dedicated to precisely segmenting text into individual sentences.

  • UniversalSplitter –

    Language-agnostic sentence boundary detector using regex patterns.

Functions:

  • detect_top_language –

    Detect the top language of the given text using py3langid.

  • log_info –

    Log an info message if verbose is enabled.

  • read_text_file –

    Read text file with automatic encoding detection.

  • validate_input –

    A decorator that validates function inputs and outputs

SentenceSplitter

SentenceSplitter(lang: str, verbose: bool = False)

A robust and versatile utility dedicated to precisely segmenting text into individual sentences.

Key Features: - Multilingual Support: Leverages language-specific algorithms and detection for broad coverage. - Fallback Mechanism: Employs a universal rule-based splitter for unsupported languages. - Intelligent Post-processing: Cleans up split sentences by filtering empty strings and rejoining stray punctuation.

Initializes the SentenceSplitter.

Parameters:

  • lang

    (str) –

    Language code (e.g., 'en', 'fr', 'auto').

  • verbose

    (bool, default: False ) –

    If True, enables verbose logging for debugging and informational messages.

Methods:

  • split_file –

    Read and split a file into sentences.

  • split_text –

    Splits a given text into a list of sentences.

Source code in src/chunklet/sentence_splitter/sentence_splitter.py
@validate_input
def __init__(
    self,
    lang: str,
    verbose: bool = False,
):
    """
    Initializes the SentenceSplitter.

    Args:
        lang: Language code (e.g., 'en', 'fr', 'auto').
        verbose: If True, enables verbose logging for debugging and informational messages.
    """
    self.verbose = verbose
    self.fallback_splitter = UniversalSplitter()
    self.lang = lang

    # Tracked to reduce log spamming about language detection
    self._last_lang_used = None

split_file

split_file(path: str | Path) -> list[str]

Read and split a file into sentences.

Parameters:

  • path

    (str | Path) –

    Path to the file to read.

Returns:

  • list[str] –

    A list of sentences extracted from the file.

Source code in src/chunklet/sentence_splitter/sentence_splitter.py
def split_file(self, path: str | Path) -> list[str]:
    """
    Read and split a file into sentences.

    Args:
        path: Path to the file to read.

    Returns:
        A list of sentences extracted from the file.
    """
    content = read_text_file(path)
    return self.split_text(content)

split_text

split_text(text: str) -> list[str]

Splits a given text into a list of sentences.

Parameters:

  • text

    (str) –

    The input text to be split.

Returns:

  • list[str] –

    A list of sentences.

Examples:

>>> splitter = SentenceSplitter(lang="en")
>>> splitter.split_text("Hello world. How are you?")
['Hello world.', 'How are you?']
>>> splitter = SentenceSplitter(lang="fr")
>>> splitter.split_text("Bonjour le monde. Comment allez-vous?")
['Bonjour le monde.', 'Comment allez-vous?']
>>> splitter = SentenceSplitter(lang="auto")
>>> splitter.split_text("Hello world. How are you?")
['Hello world.', 'How are you?']
Source code in src/chunklet/sentence_splitter/sentence_splitter.py
@validate_input
def split_text(self, text: str) -> list[str]:
    """
    Splits a given text into a list of sentences.

    Args:
        text: The input text to be split.

    Returns:
        A list of sentences.

    Examples:
        >>> splitter = SentenceSplitter(lang="en")
        >>> splitter.split_text("Hello world. How are you?")
        ['Hello world.', 'How are you?']
        >>> splitter = SentenceSplitter(lang="fr")
        >>> splitter.split_text("Bonjour le monde. Comment allez-vous?")
        ['Bonjour le monde.', 'Comment allez-vous?']
        >>> splitter = SentenceSplitter(lang="auto")
        >>> splitter.split_text("Hello world. How are you?")
        ['Hello world.', 'How are you?']
    """
    lang = self.lang

    if not text:
        log_info(self.verbose, "Input text is empty. Returning empty list.")
        return []

    if lang == "auto":
        lang_detected, confidence = detect_top_language(text)
        log_info(
            self.verbose,
            "Language detection: '{}' with confidence {}.",
            lang_detected,
            f"{round(confidence) * 10}/10",
        )
        lang = lang_detected if confidence >= 0.7 else "fallback"

    self._last_lang_used = lang

    sentences = None
    if (
        lang != "fallback"
        and (handler := self._get_lang_handler(lang, self.verbose)) is not None
    ):
        sentences = handler(text)

    # If no handler found, use fallback
    if sentences is None:
        logger.warning(
            "Using a universal rule-based splitter.\n"
            "Reason: Language not supported or detected with low confidence."
        )
        sentences = self.fallback_splitter.split(text)

    cleaned_sentences = self._clean_sentences(sentences)
    log_info(
        self.verbose,
        "Text splitted into sentences. Total sentences detected: {}",
        len(cleaned_sentences),
    )
    return cleaned_sentences

UniversalSplitter

Language-agnostic sentence boundary detector using regex patterns.

A universal splitter using Unicode-aware regex patterns for any language.

Handles
  • Unicode sentence terminators
  • Numbered lists and headings
  • Quoted sentences
  • Line breaks and whitespace
Use cases
  • Primary splitter for languages without dedicated support
  • Fallback when language-specific splitters unavailable

Methods:

  • split –

    Splits text into sentences using rule-based regex patterns.

split

split(text: str) -> list[str]

Splits text into sentences using rule-based regex patterns.

Parameters:

  • text

    (str) –

    The input text to be segmented into sentences.

Returns:

  • list[str] –

    A list of sentences after segmentation.

Source code in src/chunklet/sentence_splitter/_universal_splitter.py
def split(self, text: str) -> list[str]:
    """
    Splits text into sentences using rule-based regex patterns.

    Args:
        text: The input text to be segmented into sentences.

    Returns:
        A list of sentences after segmentation.
    """

    def mask(match: re.Match, norm_map: dict):
        # Generate the integer hash and Convert to string
        # because re.sub MUST return a string
        # Also fence them for easy detection
        hashed_str = f"##{hash(match.group())}##"

        # Store the mapping for later reconstruction
        norm_map[hashed_str] = match.group()
        return hashed_str

    def unmask(match: re.Match, norm_map: dict):
        return norm_map.get(match.group(), match.group())

    text = FLATTENED_NUMBERED_LIST_PATTERN.sub(r"\n \1", text.strip())

    # Normalize to protect them
    norm_map = {}
    text = QUOTE_OR_PAREN_PATTERN.sub(lambda m: mask(m, norm_map), text)
    text = NUMBERED_LIST_PATTERN.sub(lambda m: mask(m, norm_map), text)

    # Firstly, split base on punctuation
    # then split further on newline
    final_sentences = []
    sentences = SENTENCE_END_PATTERN.split(text.strip())

    for sent in sentences:
        if sent:
            final_sentences.extend(sent.strip().splitlines())

    # Restore the normalization
    return [
        HASHED_PATTERN.sub(lambda m: unmask(m, norm_map), sent)
        for sent in final_sentences
        if sent.strip()
    ]

detect_top_language

detect_top_language(text: str) -> tuple[str, float]

Detect the top language of the given text using py3langid.

The LanguageIdentifier is built lazily on first use and cached for reuse.

Parameters:

  • text

    (str) –

    The input text to detect the language for.

Returns:

  • str –

    A tuple of the ISO 639-1 language code and its confidence in [0, 1].

  • float –

    Confidence depends on the py3langid model, so treat it as

  • tuple[str, float] –

    approximate rather than a threshold you can rely on across versions.

Raises:

  • ImportError –

    If py3langid is not installed.

Examples:

>>> lang, confidence = detect_top_language("This sentence is written in English.")
>>> lang, confidence > 0.8
('en', True)
>>> detect_top_language("Ceci est une phrase ecrite en francais.")[0]
'fr'
>>> code, confidence = detect_top_language("")
>>> round(confidence, 2)
0.01
Source code in src/chunklet/common/lang_detection.py
def detect_top_language(text: str) -> tuple[str, float]:
    """Detect the top language of the given text using py3langid.

    The LanguageIdentifier is built lazily on first use and cached for reuse.

    Args:
        text: The input text to detect the language for.

    Returns:
        A tuple of the ISO 639-1 language code and its confidence in ``[0, 1]``.
        Confidence depends on the ``py3langid`` model, so treat it as
        approximate rather than a threshold you can rely on across versions.

    Raises:
        ImportError: If py3langid is not installed.

    Examples:
        >>> lang, confidence = detect_top_language("This sentence is written in English.")
        >>> lang, confidence > 0.8
        ('en', True)
        >>> detect_top_language("Ceci est une phrase ecrite en francais.")[0]
        'fr'
        >>> code, confidence = detect_top_language("")
        >>> round(confidence, 2)
        0.01
    """
    global _lang_identifier
    if _lang_identifier is None:
        try:
            from py3langid.langid import MODEL_FILE, LanguageIdentifier

            _lang_identifier = LanguageIdentifier.from_model_file(
                MODEL_FILE, norm_probs=True
            )
        except ImportError as e:  # pragma: no cover
            raise ImportError(
                "The 'py3langid' library is required for auto language detection. "
                "Please install it with 'pip install 'py3langid>=0.4.0,<0.5.0'' "
                "or install the lang-detect extra with 'pip install 'chunklet-py[lang-detect]''"
            ) from e

    return _lang_identifier.classify(text)

log_info

log_info(verbose: bool, *args, **kwargs) -> None

Log an info message if verbose is enabled.

This is a convenience function that only logs when verbose mode is enabled, avoiding unnecessary log output in production.

Parameters:

  • verbose

    (bool) –

    If True, logs the message; if False, does nothing.

  • *args

    –

    Positional arguments passed to logger.info().

  • **kwargs

    –

    Keyword arguments passed to logger.info().

Example

log_info(True, "Processing file: {}", filepath) Processing file: /path/to/file log_info(False, "This will not be logged") (no output)

Source code in src/chunklet/common/logging_utils.py
def log_info(verbose: bool, *args, **kwargs) -> None:
    """Log an info message if verbose is enabled.

    This is a convenience function that only logs when verbose mode is enabled,
    avoiding unnecessary log output in production.

    Args:
        verbose: If True, logs the message; if False, does nothing.
        *args: Positional arguments passed to logger.info().
        **kwargs: Keyword arguments passed to logger.info().

    Example:
        >>> log_info(True, "Processing file: {}", filepath)
        Processing file: /path/to/file
        >>> log_info(False, "This will not be logged")
        (no output)
    """
    if verbose:
        logger.info(*args, **kwargs)

read_text_file

read_text_file(path: str | Path) -> str

Read text file with automatic encoding detection.

Parameters:

  • path

    (str | Path) –

    File path to read.

Returns:

  • str –

    File content.

Raises:

Source code in src/chunklet/common/path_utils.py
@validate_input
def read_text_file(path: str | Path) -> str:
    """Read text file with automatic encoding detection.

    Args:
        path: File path to read.

    Returns:
        File content.

    Raises:
        FileProcessingError: If file cannot be read.
    """
    from charset_normalizer import from_path

    path = Path(path)

    if not path.exists():
        raise FileProcessingError(f"File does not exist: {path}")

    if is_binary_file(path):
        raise FileProcessingError(f"Binary file not supported: {path}")

    match = from_path(str(path)).best()
    return str(match) if match else ""

validate_input

validate_input(fn)

A decorator that validates function inputs and outputs

A wrapper around Pydantic's validate_call that catchesValidationError and re-raises it as a more user-friendly InvalidInputError.

Source code in src/chunklet/common/validation.py
def validate_input(fn):
    """
    A decorator that validates function inputs and outputs

    A wrapper around Pydantic's `validate_call` that catches`ValidationError` and re-raises it as a more user-friendly `InvalidInputError`.
    """
    validated_fn = validate_call(fn, config=ConfigDict(arbitrary_types_allowed=True))

    @wraps(fn)
    def wrapper(*args, **kwargs):
        try:
            return validated_fn(*args, **kwargs)
        except ValidationError as e:
            raise InvalidInputError(pretty_errors(e)) from None

    return wrapper