Language-agnostic sentence boundary detector using regex patterns.
A universal splitter using Unicode-aware regex patterns for any language.
Handles
- Unicode sentence terminators
- Numbered lists and headings
- Quoted sentences
- Line breaks and whitespace
Use cases
- Primary splitter for languages without dedicated support
- Fallback when language-specific splitters unavailable
Methods:
-
split
–
Splits text into sentences using rule-based regex patterns.
split
split(text: str) -> list[str]
Splits text into sentences using rule-based regex patterns.
Parameters:
-
text
(str)
–
The input text to be segmented into sentences.
Returns:
-
list[str]
–
A list of sentences after segmentation.
Source code in src/chunklet/sentence_splitter/_universal_splitter.py
| def split(self, text: str) -> list[str]:
"""
Splits text into sentences using rule-based regex patterns.
Args:
text: The input text to be segmented into sentences.
Returns:
A list of sentences after segmentation.
"""
def mask(match: re.Match, norm_map: dict):
# Generate the integer hash and Convert to string
# because re.sub MUST return a string
# Also fence them for easy detection
hashed_str = f"##{hash(match.group())}##"
# Store the mapping for later reconstruction
norm_map[hashed_str] = match.group()
return hashed_str
def unmask(match: re.Match, norm_map: dict):
return norm_map.get(match.group(), match.group())
text = FLATTENED_NUMBERED_LIST_PATTERN.sub(r"\n \1", text.strip())
# Normalize to protect them
norm_map = {}
text = QUOTE_OR_PAREN_PATTERN.sub(lambda m: mask(m, norm_map), text)
text = NUMBERED_LIST_PATTERN.sub(lambda m: mask(m, norm_map), text)
# Firstly, split base on punctuation
# then split further on newline
final_sentences = []
sentences = SENTENCE_END_PATTERN.split(text.strip())
for sent in sentences:
if sent:
final_sentences.extend(sent.strip().splitlines())
# Restore the normalization
return [
HASHED_PATTERN.sub(lambda m: unmask(m, norm_map), sent)
for sent in final_sentences
if sent.strip()
]
|