Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Recommenders team Recommenders ImplicitCF

From Leeroopedia
Revision as of 16:28, 16 February 2026 by Admin (talk | contribs) (Auto-imported from implementations/Recommenders_team_Recommenders_ImplicitCF.md)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)


Knowledge Sources
Domains Recommendation Systems, Graph Neural Networks, Data Processing
Last Updated 2026-02-10 00:00 GMT

Overview

ImplicitCF is a data processing class for Graph Convolutional Network (GCN) based recommendation models that use implicit feedback, such as LightGCN.

Description

The ImplicitCF class handles all data preprocessing required by graph-based collaborative filtering models. It takes train and optional test DataFrames, reindexes user and item IDs to contiguous integers starting from zero, and builds a sparse user-item interaction matrix. The class constructs a normalized adjacency matrix for graph convolution by forming a bipartite graph structure [0, R; R^T, 0] and applying symmetric normalization (D^{-1/2} A D^{-1/2}), where D is the degree matrix and A is the adjacency matrix. Only interactions with ratings greater than zero are retained, consistent with the implicit feedback paradigm.

The train_loader method provides mini-batch sampling with BPR-style triplets (user, positive item, negative item). For each user in a batch, one positive item is randomly sampled from the user's interaction history, and one negative item is randomly sampled from items the user has not interacted with.

The adjacency matrix can be cached to disk to avoid recomputation, which is controlled by the optional adj_dir parameter.

Usage

Use ImplicitCF when preparing data for GCN-based recommendation models such as LightGCN. It is the required data component for any model that needs a normalized adjacency matrix built from user-item implicit interactions, along with BPR-style negative sampling for training.

Code Reference

Source Location

Signature

class ImplicitCF(object):
    def __init__(
        self,
        train,
        test=None,
        adj_dir=None,
        col_user=DEFAULT_USER_COL,
        col_item=DEFAULT_ITEM_COL,
        col_rating=DEFAULT_RATING_COL,
        col_prediction=DEFAULT_PREDICTION_COL,
        seed=None,
    ):

    def get_norm_adj_mat(self):
        # Returns: scipy.sparse.csr_matrix

    def create_norm_adj_mat(self):
        # Returns: scipy.sparse.csr_matrix

    def train_loader(self, batch_size):
        # Returns: (numpy.ndarray, numpy.ndarray, numpy.ndarray)

Import

from recommenders.models.deeprec.DataModel.ImplicitCF import ImplicitCF

I/O Contract

Inputs

Name Type Required Description
train pandas.DataFrame Yes Training data with at least columns (col_user, col_item, col_rating)
test pandas.DataFrame No Test data with at least columns (col_user, col_item, col_rating). If None, only training data is processed
adj_dir str No Directory to save/load adjacency matrices. If None, matrices are created but not saved to disk
col_user str No User column name (default: DEFAULT_USER_COL)
col_item str No Item column name (default: DEFAULT_ITEM_COL)
col_rating str No Rating column name (default: DEFAULT_RATING_COL)
col_prediction str No Prediction column name (default: DEFAULT_PREDICTION_COL)
seed int No Random seed for reproducibility

Outputs

Name Type Description
get_norm_adj_mat() scipy.sparse.csr_matrix Normalized adjacency matrix of shape (n_users + n_items, n_users + n_items) with symmetric normalization
train_loader(batch_size) (numpy.ndarray, numpy.ndarray, numpy.ndarray) Tuple of (sampled_users, positive_items, negative_items) arrays for BPR-style training

Key Attributes

Attribute Type Description
n_users int Total number of unique users across train and test sets
n_items int Total number of unique items across train and test sets
n_users_in_train int Number of unique users in the training set only
R scipy.sparse.dok_matrix User-item interaction matrix of shape (n_users, n_items)
interact_status pandas.DataFrame DataFrame recording the set of items each user has interacted with
user2id dict Mapping from original user IDs to reindexed integer IDs
id2user dict Mapping from reindexed integer IDs back to original user IDs
item2id dict Mapping from original item IDs to reindexed integer IDs
id2item dict Mapping from reindexed integer IDs back to original item IDs

Usage Examples

Basic Usage

import pandas as pd
from recommenders.models.deeprec.DataModel.ImplicitCF import ImplicitCF

# Prepare training data
train_df = pd.DataFrame({
    "userID": [0, 0, 1, 1, 2],
    "itemID": [0, 1, 1, 2, 0],
    "rating": [1, 1, 1, 1, 1],
})

test_df = pd.DataFrame({
    "userID": [0, 1, 2],
    "itemID": [2, 0, 1],
    "rating": [1, 1, 1],
})

# Initialize the data model
data = ImplicitCF(train=train_df, test=test_df, seed=42)

# Get the normalized adjacency matrix for GCN
norm_adj_mat = data.get_norm_adj_mat()
print(f"Adjacency matrix shape: {norm_adj_mat.shape}")
# Output: Adjacency matrix shape: (n_users + n_items, n_users + n_items)

# Sample a training batch with BPR triplets
batch_size = 2
users, pos_items, neg_items = data.train_loader(batch_size)
print(f"Users: {users}, Pos items: {pos_items}, Neg items: {neg_items}")

With Adjacency Matrix Caching

from recommenders.models.deeprec.DataModel.ImplicitCF import ImplicitCF

# Use adj_dir to cache the adjacency matrix on disk
data = ImplicitCF(
    train=train_df,
    test=test_df,
    adj_dir="/tmp/adj_cache",
    seed=42,
)

# First call creates and saves the matrix; subsequent calls load from disk
norm_adj = data.get_norm_adj_mat()

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment