Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Explodinggradients Ragas LLM Benchmarking
- Workflow:LLMBook zh LLMBook zh github io LLM Pretraining
- Workflow:Run llama Llama index RAG Query Pipeline
- Workflow:Nautechsystems Nautilus trader Backtest with BacktestNode
- Workflow:Lucidrains X transformers Autoregressive Language Modeling
- Workflow:Huggingface Diffusers Model Quantization
- Workflow:Astronomer Astronomer cosmos Local dbt DAG rendering
- Workflow:Microsoft DeepSpeedExamples CIFAR10 Getting Started
- Workflow:Mlc ai Mlc llm REST API Serving
- Workflow:Online ml River Drift Adaptive Classification
Principles
- Principle:PeterL1n BackgroundMattingV2 Video output writing
- Principle:Apache Flink Source Chain Composition
- Principle:DevExpress Testcafe Configuration Loading
- Principle:ClickHouse ClickHouse Unaligned Memory Access
- Principle:Langgenius Dify Application Factory
- Principle:Allenai Open instruct Reward Extraction
- Principle:Facebookresearch Habitat lab HRL Evaluation
- Principle:DistrictDataLabs Yellowbrick Type Introspection
- Principle:Huggingface Transformers Fully Sharded Data Parallelism
- Principle:Speechbrain Speechbrain Whisper Finetuning With LR Scheduling
Implementations
- Implementation:Apache Airflow Provider Yaml Schema
- Implementation:CARLA simulator Carla Navigation
- Implementation:Explodinggradients Ragas ContextRecall Metric
- Implementation:Microsoft Onnxruntime CPU Scale
- Implementation:Google deepmind Mujoco mj saveLastXML
- Implementation:Interpretml Interpret PartitionMultiDimensionalCorner
- Implementation:Online ml River Base MultiOutput
- Implementation:Webdriverio Webdriverio MSPO Overwrite
- Implementation:Eventual Inc Daft Catalog From Iceberg
- Implementation:Ucbepic Docetl PipelineVisualization
Heuristics
- Heuristic:TobikoData Sqlmesh Model Change Categorization
- Heuristic:Openai Whisper KV Cache Optimization
- Heuristic:Huggingface Alignment handbook Sequence Packing Strategy
- Heuristic:Scikit learn contrib Imbalanced learn Sampling Before Split Leakage
- Heuristic:Openai CLIP Class Name Curation
- Heuristic:Mlfoundations Open flamingo Gradient Clipping Max Norm
- Heuristic:DataTalksClub Data engineering zoomcamp CSV Chunk Size Optimization
- Heuristic:Hpcaitech ColossalAI Gradient Checkpointing Memory Tip
- Heuristic:Neuml Txtai MacOS Stability Workarounds
- Heuristic:Triton inference server Server Documentation Standards
Environments
- Environment:Cohere ai Cohere python Python SDK Runtime
- Environment:BerriAI Litellm Provider API Credentials
- Environment:Marker Inc Korea AutoRAG Japanese NLP Dependencies
- Environment:Ggml org Ggml Metal GPU Environment
- Environment:Apache Kafka Release Toolchain Environment
- Environment:Princeton nlp Tree of thought llm Python OpenAI
- Environment:Sgl project Sglang Performance Dashboard
- Environment:ArroyoSystems Arroyo Python UDF Runtime
- Environment:Run llama Llama index OpenAI API Configuration
- Environment:Lm sys FastChat GPU CUDA Inference