Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Lm sys FastChat LoRA QLoRA Finetuning
- Workflow:ContextualAI HALOs Model Evaluation
- Workflow:Recommenders team Recommenders ALS Spark Recommendation
- Workflow:Huggingface Datatrove FineWeb Dataset Creation
- Workflow:NVIDIA NeMo Aligner DPO Training
- Workflow:Huggingface Datasets Format Conversion
- Workflow:Openai Openai node Structured Output Parsing
- Workflow:OpenRLHF OpenRLHF Math Reasoning Training
- Workflow:Langgenius Dify Plugin Installation and Configuration
- Workflow:Eric mitchell Direct preference optimization Custom Dataset Integration
Principles
- Principle:Online ml River Streaming Data Loading
- Principle:Lance format Lance Tag Management
- Principle:Haifengl Smile Nearest Neighbor Query
- Principle:Dotnet Machinelearning Latent Dirichlet Allocation
- Principle:Mit han lab Llm awq LM Evaluation Harness Adaptation
- Principle:Cohere ai Cohere python Text Embedding
- Principle:Apache Kafka SVN Artifact Staging
- Principle:Helicone Helicone Legal Compliance
- Principle:Cleanlab Cleanlab Coteaching Algorithm
- Principle:Interpretml Interpret Bagged Gradient Boosting
Implementations
- Implementation:Speechbrain Speechbrain CVSS Extract Code
- Implementation:Google deepmind Mujoco mjr makeContext
- Implementation:Deepspeedai DeepSpeed Reduction Utils
- Implementation:ThreeSR Awesome Inference Time Scaling Get Paper Info Function
- Implementation:InternLM Lmdeploy Tensor
- Implementation:Openai Openai python Moderation Create Params
- Implementation:FMInference FlexLLMGen DeepSpeed GEMM Test
- Implementation:Haifengl Smile NearestNeighborGraph
- Implementation:SeleniumHQ Selenium Closure SafeStyleSheet
- Implementation:Pyro ppl Pyro Settings
Heuristics
- Heuristic:Mlfoundations Open flamingo Deterministic Shard Shuffling
- Heuristic:Intel Ipex llm CCL Distributed Training Tips
- Heuristic:Openai Whisper No Speech Detection
- Heuristic:Online ml River ARF Drift Detection Sensitivity
- Heuristic:Cohere ai Cohere python ToolCallV2 Auto UUID Override
- Heuristic:Promptfoo Promptfoo Warning Deprecated Cache Migration
- Heuristic:Spcl Graph of thoughts Four Bit Quantization For Local LLMs
- Heuristic:Mbzuai oryx Awesome LLM Post training Checkpoint Every 3 Papers
- Heuristic:Apache Hudi Record Level Index Optimization
- Heuristic:Volcengine Verl Sequence Length Balancing
Environments
- Environment:Google deepmind Mujoco MJX JAX Environment
- Environment:Farama Foundation Gymnasium Python 3 10 Runtime
- Environment:FMInference FlexLLMGen NVMe Disk
- Environment:ClickHouse ClickHouse Python3 Test Environment
- Environment:DataTalksClub Data engineering zoomcamp Dlt BigQuery Environment
- Environment:Apache Druid Web Console Development
- Environment:Speechbrain Speechbrain HuggingFace Transformers
- Environment:Mit han lab Llm awq CUDA Build Environment
- Environment:Princeton nlp SimPO CUDA Training
- Environment:Datahub project Datahub Java Build