Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Evidentlyai Evidently Legacy Sentence Count Feature

From Leeroopedia
Knowledge Sources
Domains ML Monitoring, Text Analysis, Data Quality
Last Updated 2026-02-14 12:00 GMT

Overview

Provides a generated feature that counts the number of sentences in each value of a specified text column using a regular expression-based sentence splitter.

Description

The SentenceCount class extends ApplyColumnGeneratedFeature to compute a numerical feature representing the number of sentences in a text value. Sentence splitting is performed using a compiled regular expression pattern stored as a class variable _reg:

re.compile(r"(?<!\w\.\w.)(?<![A-Z][a-z]\.)(?<=\.|\?)\s")

This pattern splits text on whitespace that follows a period or question mark, while avoiding splits after common abbreviations (e.g., "Dr.", "Mr.") and decimal numbers (e.g., "3.14"). The sentence count is computed as the number of segments produced by the split, with a minimum return value of 1 (ensuring that non-empty text always counts as at least one sentence).

Special handling is provided for None values and NaN floats, which return 0. The feature type is ColumnType.Numerical.

The class uses a display_name_template class variable set to "Sentence Count for {column_name}" for automatic display naming.

Usage

Use this feature to monitor text length and complexity characteristics. It is useful for tracking the verbosity of LLM outputs, detecting anomalies in text generation, or measuring data drift in text-based features.

Code Reference

Source Location

Signature

class SentenceCount(ApplyColumnGeneratedFeature):
    class Config:
        type_alias = "evidently:feature:SentenceCount"

    __feature_type__: ClassVar = ColumnType.Numerical
    _reg: ClassVar[re.Pattern] = re.compile(r"(?<!\w\.\w.)(?<![A-Z][a-z]\.)(?<=\.|\?)\s")
    display_name_template: ClassVar = "Sentence Count for {column_name}"
    column_name: str

    def __init__(self, column_name: str, display_name: Optional[str] = None): ...
    def apply(self, value: Any): ...

Import

from evidently.legacy.features.sentence_count_feature import SentenceCount

I/O Contract

Inputs

Name Type Required Description
column_name str Yes Name of the text column in the DataFrame to analyze
display_name Optional[str] No Custom display name for the feature (defaults to template-based name)

Outputs

Name Type Description
return int (per row) Number of sentences in the text value; 0 for None/NaN values, minimum 1 for non-empty text

Usage Examples

from evidently.legacy.features.sentence_count_feature import SentenceCount

# Create the feature for a column named "text"
sentence_feature = SentenceCount(column_name="text")

# With a custom display name
sentence_feature = SentenceCount(
    column_name="response",
    display_name="Response Sentence Count"
)

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment