Implementation:Recommenders team Recommenders ImplicitCF
| Knowledge Sources | |
|---|---|
| Domains | Recommendation Systems, Graph Neural Networks, Data Processing |
| Last Updated | 2026-02-10 00:00 GMT |
Overview
ImplicitCF is a data processing class for Graph Convolutional Network (GCN) based recommendation models that use implicit feedback, such as LightGCN.
Description
The ImplicitCF class handles all data preprocessing required by graph-based collaborative filtering models. It takes train and optional test DataFrames, reindexes user and item IDs to contiguous integers starting from zero, and builds a sparse user-item interaction matrix. The class constructs a normalized adjacency matrix for graph convolution by forming a bipartite graph structure [0, R; R^T, 0] and applying symmetric normalization (D^{-1/2} A D^{-1/2}), where D is the degree matrix and A is the adjacency matrix. Only interactions with ratings greater than zero are retained, consistent with the implicit feedback paradigm.
The train_loader method provides mini-batch sampling with BPR-style triplets (user, positive item, negative item). For each user in a batch, one positive item is randomly sampled from the user's interaction history, and one negative item is randomly sampled from items the user has not interacted with.
The adjacency matrix can be cached to disk to avoid recomputation, which is controlled by the optional adj_dir parameter.
Usage
Use ImplicitCF when preparing data for GCN-based recommendation models such as LightGCN. It is the required data component for any model that needs a normalized adjacency matrix built from user-item implicit interactions, along with BPR-style negative sampling for training.
Code Reference
Source Location
- Repository: Recommenders
- File: recommenders/models/deeprec/DataModel/ImplicitCF.py
- Lines: 1-230
Signature
class ImplicitCF(object):
def __init__(
self,
train,
test=None,
adj_dir=None,
col_user=DEFAULT_USER_COL,
col_item=DEFAULT_ITEM_COL,
col_rating=DEFAULT_RATING_COL,
col_prediction=DEFAULT_PREDICTION_COL,
seed=None,
):
def get_norm_adj_mat(self):
# Returns: scipy.sparse.csr_matrix
def create_norm_adj_mat(self):
# Returns: scipy.sparse.csr_matrix
def train_loader(self, batch_size):
# Returns: (numpy.ndarray, numpy.ndarray, numpy.ndarray)
Import
from recommenders.models.deeprec.DataModel.ImplicitCF import ImplicitCF
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| train | pandas.DataFrame | Yes | Training data with at least columns (col_user, col_item, col_rating) |
| test | pandas.DataFrame | No | Test data with at least columns (col_user, col_item, col_rating). If None, only training data is processed |
| adj_dir | str | No | Directory to save/load adjacency matrices. If None, matrices are created but not saved to disk |
| col_user | str | No | User column name (default: DEFAULT_USER_COL) |
| col_item | str | No | Item column name (default: DEFAULT_ITEM_COL) |
| col_rating | str | No | Rating column name (default: DEFAULT_RATING_COL) |
| col_prediction | str | No | Prediction column name (default: DEFAULT_PREDICTION_COL) |
| seed | int | No | Random seed for reproducibility |
Outputs
| Name | Type | Description |
|---|---|---|
| get_norm_adj_mat() | scipy.sparse.csr_matrix | Normalized adjacency matrix of shape (n_users + n_items, n_users + n_items) with symmetric normalization |
| train_loader(batch_size) | (numpy.ndarray, numpy.ndarray, numpy.ndarray) | Tuple of (sampled_users, positive_items, negative_items) arrays for BPR-style training |
Key Attributes
| Attribute | Type | Description |
|---|---|---|
| n_users | int | Total number of unique users across train and test sets |
| n_items | int | Total number of unique items across train and test sets |
| n_users_in_train | int | Number of unique users in the training set only |
| R | scipy.sparse.dok_matrix | User-item interaction matrix of shape (n_users, n_items) |
| interact_status | pandas.DataFrame | DataFrame recording the set of items each user has interacted with |
| user2id | dict | Mapping from original user IDs to reindexed integer IDs |
| id2user | dict | Mapping from reindexed integer IDs back to original user IDs |
| item2id | dict | Mapping from original item IDs to reindexed integer IDs |
| id2item | dict | Mapping from reindexed integer IDs back to original item IDs |
Usage Examples
Basic Usage
import pandas as pd
from recommenders.models.deeprec.DataModel.ImplicitCF import ImplicitCF
# Prepare training data
train_df = pd.DataFrame({
"userID": [0, 0, 1, 1, 2],
"itemID": [0, 1, 1, 2, 0],
"rating": [1, 1, 1, 1, 1],
})
test_df = pd.DataFrame({
"userID": [0, 1, 2],
"itemID": [2, 0, 1],
"rating": [1, 1, 1],
})
# Initialize the data model
data = ImplicitCF(train=train_df, test=test_df, seed=42)
# Get the normalized adjacency matrix for GCN
norm_adj_mat = data.get_norm_adj_mat()
print(f"Adjacency matrix shape: {norm_adj_mat.shape}")
# Output: Adjacency matrix shape: (n_users + n_items, n_users + n_items)
# Sample a training batch with BPR triplets
batch_size = 2
users, pos_items, neg_items = data.train_loader(batch_size)
print(f"Users: {users}, Pos items: {pos_items}, Neg items: {neg_items}")
With Adjacency Matrix Caching
from recommenders.models.deeprec.DataModel.ImplicitCF import ImplicitCF
# Use adj_dir to cache the adjacency matrix on disk
data = ImplicitCF(
train=train_df,
test=test_df,
adj_dir="/tmp/adj_cache",
seed=42,
)
# First call creates and saves the matrix; subsequent calls load from disk
norm_adj = data.get_norm_adj_mat()