Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Duckdb Duckdb Building From Source
- Workflow:Deepspeedai DeepSpeed Pipeline Parallel Training
- Workflow:MaterializeInc Materialize Docker Image Build
- Workflow:Cohere ai Cohere python Tool Use Agentic Chat
- Workflow:Cohere ai Cohere python Semantic Search With Rerank
- Workflow:Evidentlyai Evidently Data Drift Monitoring
- Workflow:Speechbrain Speechbrain Speaker Embedding Training
- Workflow:Huggingface Trl PPO RLHF Training
- Workflow:Hiyouga LLaMA Factory Full Parameter SFT
- Workflow:Explodinggradients Ragas Prompt Evaluation And Iteration
Principles
- Principle:FMInference FlexLLMGen DeepSpeed Initialization
- Principle:Datajuicer Data juicer Partition Size Optimization
- Principle:Ggml org Llama cpp Speculative Generation
- Principle:HKUDS AI Trader LLM Invocation
- Principle:Kornia Kornia ONNX Sequential Pipeline
- Principle:Datahub project Datahub Docker Prerequisites Validation
- Principle:Huggingface Datasets Dataset Hub Upload
- Principle:Cohere ai Cohere python Embed Response Processing
- Principle:Kornia Kornia Grayscale Conversion
- Principle:Cleanlab Cleanlab CleanLearning Initialization
Implementations
- Implementation:Microsoft LoRA Check Copies
- Implementation:Bentoml BentoML SDK Validators
- Implementation:Mlc ai Mlc llm Preshard
- Implementation:Datajuicer Data juicer CalibrateResponseMapper
- Implementation:Mlflow Mlflow Dev Environment Setup
- Implementation:Kubeflow Pipelines Dsl Condition
- Implementation:BerriAI Litellm Cooldown Cache
- Implementation:FlowiseAI Flowise ItemCard
- Implementation:ARISE Initiative Robosuite TrajUtils
- Implementation:SeleniumHQ Selenium Closure Dom Forms
Heuristics
- Heuristic:Haifengl Smile Quarkus Async Context Handling
- Heuristic:Huggingface Open r1 Test Batch Early Termination
- Heuristic:Huggingface Trl QLoRA BF16 Adapter Casting
- Heuristic:Facebookresearch Habitat lab Mini Batch Environment Divisibility
- Heuristic:NVIDIA NeMo Curator GPU Memory Resource Allocation
- Heuristic:Princeton nlp Tree of thought llm Global State Token Counting
- Heuristic:Dagster io Dagster Retry Strategy Configuration
- Heuristic:Trailofbits Fickling Injection Mode Selection
- Heuristic:Gretelai Gretel synthetics Binary Encoder Cutoff
- Heuristic:Datajuicer Data juicer Operator Fusion Rules
Environments
- Environment:Turboderp org Exllamav2 Flash Attention Backend
- Environment:Hpcaitech ColossalAI ColossalChat Training Environment
- Environment:Huggingface Diffusers Quantization Environment
- Environment:ArroyoSystems Arroyo Kubernetes Deployment
- Environment:Apache Kafka JVM Runtime Environment
- Environment:Lance format Lance SIMD And Platform Requirements
- Environment:InternLM Lmdeploy Python Dependencies
- Environment:Spcl Graph of thoughts OpenAI API Access
- Environment:Sgl project Sglang CUDA
- Environment:Unstructured IO Unstructured All Docs