Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Spotify Luigi Database Ingestion Pipeline
- Workflow:NVIDIA NeMo Aligner Reward Model Training
- Workflow:Farama Foundation Gymnasium Vectorized Environment Training
- Workflow:DataExpert io Data engineer handbook PySpark Job Testing
- Workflow:SeleniumHQ Selenium Page Object Pattern Testing
- Workflow:Heibaiying BigData Notes Hadoop MapReduce Word Count
- Workflow:Kornia Kornia Image Feature Matching
- Workflow:Diagram of thought Diagram of thought DoT Trace Extraction
- Workflow:Dotnet Machinelearning Time Series Forecasting
- Workflow:SeleniumHQ Selenium Contributor Development Workflow
Principles
- Principle:Scikit learn Scikit learn Dataset Loading
- Principle:Arize ai Phoenix Prompt Template Rendering
- Principle:Datajuicer Data juicer Data Mapping Transformation
- Principle:Haotian liu LLaVA Benchmark Metric Computation
- Principle:Microsoft BIPIA Training Data Tokenization
- Principle:FlagOpen FlagEmbedding Contrastive Embedding Training
- Principle:Microsoft Autogen Team Deployment
- Principle:Cleanlab Cleanlab Object Detection Issue Filtering
- Principle:Bitsandbytes foundation Bitsandbytes CPU SIMD Dequantization
- Principle:Explodinggradients Ragas LLM as Judge Metric
Implementations
- Implementation:Mit han lab Llm awq LLaVA Image Processing
- Implementation:Ollama Ollama Convert GptOss
- Implementation:Datajuicer Data juicer VideoCameraCalibrationStaticMogeMapper
- Implementation:Avhz RustQuant Newton Raphson
- Implementation:Hpcaitech ColossalAI DatasetEvaluator
- Implementation:DistrictDataLabs Yellowbrick Rank2D Visualizer
- Implementation:Datahub project Datahub FilterTransformer Transform
- Implementation:MaterializeInc Materialize CLI Optbench
- Implementation:CARLA simulator Carla BasicAgent Done
- Implementation:Onnx Onnx ReferenceEvaluator Init
Heuristics
- Heuristic:Pytorch Serve Ampere Tensor Core Optimization
- Heuristic:Openai CLIP Class Name Curation
- Heuristic:NVIDIA TransformerEngine Build Optimization Tips
- Heuristic:OpenGVLab InternVL LoRA Alpha Scaling
- Heuristic:LaurentMazare Tch rs CuDNN Benchmark Mode
- Heuristic:Microsoft Onnxruntime Convergence Debugging Tips
- Heuristic:Huggingface Transformers FSDP Activation Checkpointing Tip
- Heuristic:Gretelai Gretel synthetics Memory Chunking For Normalization
- Heuristic:Pola rs Polars GPU Aggregation Join Speedup
- Heuristic:Scikit learn contrib Imbalanced learn Sampling Before Split Leakage
Environments
- Environment:Axolotl ai cloud Axolotl Python Runtime
- Environment:Anthropics Anthropic sdk python Python SDK Core Environment
- Environment:LLMBook zh LLMBook zh github io VLLM Inference Environment
- Environment:NVIDIA NeMo Curator NVIDIA DALI
- Environment:Sgl project Sglang CUDA SM100
- Environment:Helicone Helicone Python ClickHouse Migrations
- Environment:DataTalksClub Data engineering zoomcamp Kestra Orchestration Environment
- Environment:Googleapis Python genai Vertex AI Service Account
- Environment:Neuml Txtai Docker Deployment Environment
- Environment:Microsoft Agent framework Python 3 10 Runtime