Implementation:Evidentlyai Evidently Legacy Sentence Count Feature
| Knowledge Sources | |
|---|---|
| Domains | ML Monitoring, Text Analysis, Data Quality |
| Last Updated | 2026-02-14 12:00 GMT |
Overview
Provides a generated feature that counts the number of sentences in each value of a specified text column using a regular expression-based sentence splitter.
Description
The SentenceCount class extends ApplyColumnGeneratedFeature to compute a numerical feature representing the number of sentences in a text value. Sentence splitting is performed using a compiled regular expression pattern stored as a class variable _reg:
re.compile(r"(?<!\w\.\w.)(?<![A-Z][a-z]\.)(?<=\.|\?)\s")
This pattern splits text on whitespace that follows a period or question mark, while avoiding splits after common abbreviations (e.g., "Dr.", "Mr.") and decimal numbers (e.g., "3.14"). The sentence count is computed as the number of segments produced by the split, with a minimum return value of 1 (ensuring that non-empty text always counts as at least one sentence).
Special handling is provided for None values and NaN floats, which return 0. The feature type is ColumnType.Numerical.
The class uses a display_name_template class variable set to "Sentence Count for {column_name}" for automatic display naming.
Usage
Use this feature to monitor text length and complexity characteristics. It is useful for tracking the verbosity of LLM outputs, detecting anomalies in text generation, or measuring data drift in text-based features.
Code Reference
Source Location
- Repository: Evidentlyai_Evidently
- File: src/evidently/legacy/features/sentence_count_feature.py
Signature
class SentenceCount(ApplyColumnGeneratedFeature):
class Config:
type_alias = "evidently:feature:SentenceCount"
__feature_type__: ClassVar = ColumnType.Numerical
_reg: ClassVar[re.Pattern] = re.compile(r"(?<!\w\.\w.)(?<![A-Z][a-z]\.)(?<=\.|\?)\s")
display_name_template: ClassVar = "Sentence Count for {column_name}"
column_name: str
def __init__(self, column_name: str, display_name: Optional[str] = None): ...
def apply(self, value: Any): ...
Import
from evidently.legacy.features.sentence_count_feature import SentenceCount
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| column_name | str | Yes | Name of the text column in the DataFrame to analyze |
| display_name | Optional[str] | No | Custom display name for the feature (defaults to template-based name) |
Outputs
| Name | Type | Description |
|---|---|---|
| return | int (per row) | Number of sentences in the text value; 0 for None/NaN values, minimum 1 for non-empty text |
Usage Examples
from evidently.legacy.features.sentence_count_feature import SentenceCount
# Create the feature for a column named "text"
sentence_feature = SentenceCount(column_name="text")
# With a custom display name
sentence_feature = SentenceCount(
column_name="response",
display_name="Response Sentence Count"
)