Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Dagster io Dagster Modal Serverless Pipeline
- Workflow:Vllm project Vllm Offline Text Generation
- Workflow:ChenghaoMou Text dedup MinHash LSH Deduplication
- Workflow:Bentoml BentoML Model Store Management
- Workflow:ArroyoSystems Arroyo Connection Setup
- Workflow:ChenghaoMou Text dedup Suffix Array Deduplication
- Workflow:CrewAIInc CrewAI Knowledge RAG Pipeline
- Workflow:Tensorflow Serving REST API Inference
- Workflow:Treeverse LakeFS Data Version Control With Branches
- Workflow:Scikit learn Scikit learn Supervised Classification
Principles
- Principle:Apache Paimon Distributed Write via Ray
- Principle:Langchain ai Langchain Document Preparation
- Principle:Langchain ai Langchain Distribution Building
- Principle:Hpcaitech ColossalAI Preference Data Preparation
- Principle:Cleanlab Cleanlab Dataset Health Analysis
- Principle:Protectai Llm guard Vault State Management
- Principle:Allenai Open instruct Training Arguments Configuration
- Principle:Huggingface Alignment handbook Supervised Finetuning
- Principle:Sktime Pytorch forecasting DataLoader Creation
- Principle:Openai Openai node Parsed API Invocation
Implementations
- Implementation:Kornia Kornia ONNXLoader Load Model
- Implementation:Zai org CogVideo T2V I2V Dataset Loader
- Implementation:Mistralai Client python EventStream Context
- Implementation:TobikoData Sqlmesh PlanChangePreview
- Implementation:OpenGVLab InternVL LLaVA Model Builder
- Implementation:Microsoft Onnxruntime CUDA Recv
- Implementation:Lance format Lance LegacyLogicalBinaryEncoding
- Implementation:Lance format Lance Java ReserveFragmentsOp
- Implementation:Facebookresearch Habitat lab VER PreemptionDecider
- Implementation:Interpretml Interpret LinearRegression And LogisticRegression
Heuristics
- Heuristic:Cleanlab Cleanlab KNN Distance Metric Selection
- Heuristic:Spotify Luigi Atomic File Writes
- Heuristic:Hpcaitech ColossalAI Empty Cache Between Phases
- Heuristic:Pola rs Polars Lazy Over Eager Preference
- Heuristic:CARLA simulator Carla Walker Spawn Vertical Offset
- Heuristic:CrewAIInc CrewAI Rate Limiting Strategy
- Heuristic:Speechbrain Speechbrain Gradient Clipping Strategy
- Heuristic:Marker Inc Korea AutoRAG Hybrid Retrieval Score Normalization
- Heuristic:Pola rs Polars Collect All For Diverging Queries
- Heuristic:Kornia Kornia Avoid Inplace Ops Compile
Environments
- Environment:FlagOpen FlagEmbedding Python PyTorch Environment
- Environment:CrewAIInc CrewAI Python Runtime Environment
- Environment:Heibaiying BigData Notes Flink 1 9 Environment
- Environment:Sdv dev SDV Python Runtime
- Environment:Intel Ipex llm Windows Environment
- Environment:Facebookresearch Habitat lab SLURM Distributed Environment
- Environment:Roboflow Rf detr Python GPU Environment
- Environment:Vllm project Vllm CPU Runtime
- Environment:Apache Spark JDK Build Environment
- Environment:Marker Inc Korea AutoRAG Vector Database Backends