Skip to content

chunklet.sentence_splitter._universal_splitter

Classes:

  • UniversalSplitter –

    Language-agnostic sentence boundary detector using regex patterns.

UniversalSplitter

Language-agnostic sentence boundary detector using regex patterns.

A universal splitter using Unicode-aware regex patterns for any language.

Handles
  • Unicode sentence terminators
  • Numbered lists and headings
  • Quoted sentences
  • Line breaks and whitespace
Use cases
  • Primary splitter for languages without dedicated support
  • Fallback when language-specific splitters unavailable

Methods:

  • split –

    Splits text into sentences using rule-based regex patterns.

split

split(text: str) -> list[str]

Splits text into sentences using rule-based regex patterns.

Parameters:

  • text

    (str) –

    The input text to be segmented into sentences.

Returns:

  • list[str] –

    A list of sentences after segmentation.

Source code in src/chunklet/sentence_splitter/_universal_splitter.py
def split(self, text: str) -> list[str]:
    """
    Splits text into sentences using rule-based regex patterns.

    Args:
        text: The input text to be segmented into sentences.

    Returns:
        A list of sentences after segmentation.
    """

    def mask(match: re.Match, norm_map: dict):
        # Generate the integer hash and Convert to string
        # because re.sub MUST return a string
        # Also fence them for easy detection
        hashed_str = f"##{hash(match.group())}##"

        # Store the mapping for later reconstruction
        norm_map[hashed_str] = match.group()
        return hashed_str

    def unmask(match: re.Match, norm_map: dict):
        return norm_map.get(match.group(), match.group())

    text = FLATTENED_NUMBERED_LIST_PATTERN.sub(r"\n \1", text.strip())

    # Normalize to protect them
    norm_map = {}
    text = QUOTE_OR_PAREN_PATTERN.sub(lambda m: mask(m, norm_map), text)
    text = NUMBERED_LIST_PATTERN.sub(lambda m: mask(m, norm_map), text)

    # Firstly, split base on punctuation
    # then split further on newline
    final_sentences = []
    sentences = SENTENCE_END_PATTERN.split(text.strip())

    for sent in sentences:
        if sent:
            final_sentences.extend(sent.strip().splitlines())

    # Restore the normalization
    return [
        HASHED_PATTERN.sub(lambda m: unmask(m, norm_map), sent)
        for sent in final_sentences
        if sent.strip()
    ]