Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Huggingface Optimum Accelerated Inference Pipeline
- Workflow:Snorkel team Snorkel Data Augmentation
- Workflow:Eventual Inc Daft SQL Query Analytics
- Workflow:Alibaba ROLL DPO Training Pipeline
- Workflow:Kserve Kserve InferenceGraph Pipeline
- Workflow:Cohere ai Cohere python Streaming Chat
- Workflow:PacktPublishing LLM Engineers Handbook Digital Data ETL
- Workflow:Obss Sahi COCO Dataset Slicing
- Workflow:Facebookresearch Habitat lab Custom Task Extension
- Workflow:Apache Spark Application Submission
Principles
- Principle:Haotian liu LLaVA Checkpoint Extraction
- Principle:DataExpert io Data engineer handbook Windowed Aggregation
- Principle:LLMBook zh LLMBook zh github io LoRA Adapter Merging
- Principle:ArroyoSystems Arroyo SQL Query Validation
- Principle:PacktPublishing LLM Engineers Handbook Pipeline Reporting
- Principle:Kserve Kserve Admission Webhook Processing
- Principle:Microsoft Playwright Identify Interception Targets
- Principle:Onnx Onnx Tensor Specification
- Principle:Shiyu coder Kronos Topk Dropout Portfolio Strategy
- Principle:Apache Kafka GPG Prerequisite Verification
Implementations
- Implementation:Webdriverio Webdriverio Bidi NodeConnection
- Implementation:Mlflow Mlflow Build Docker
- Implementation:Deepspeedai DeepSpeed IO Handle
- Implementation:Interpretml Interpret BinSumsInteraction
- Implementation:OpenGVLab InternVL Classification Utils
- Implementation:CrewAIInc CrewAI Composio Tool
- Implementation:Datahub project Datahub RestEmitterEmitMode
- Implementation:Microsoft Onnxruntime Numpy Output Extraction
- Implementation:Open compass VLMEvalKit MMMath
- Implementation:NVIDIA NeMo Curator SemanticDeduplicationWorkflow for Video
Heuristics
- Heuristic:Princeton nlp Tree of thought llm Ad Hoc Value Map Scoring
- Heuristic:Huggingface Alignment handbook DDP Bias Buffer Ignore
- Heuristic:Fastai Fastbook Weight Decay Tuning
- Heuristic:Openai Evals Model Graded Eval Design
- Heuristic:Pola rs Polars Collect All For Diverging Queries
- Heuristic:Apache Airflow Variable Access Pattern
- Heuristic:Langchain ai Langgraph Stream Mode Selection
- Heuristic:Google deepmind Dm control Tolerance Reward Tuning
- Heuristic:Avhz RustQuant Interpolation Method Selection
- Heuristic:Mlflow Mlflow Batch Logging Size Limits
Environments
- Environment:Nautechsystems Nautilus trader Databento API Credentials
- Environment:Deepspeedai DeepSpeed CUDA GPU Environment
- Environment:DistrictDataLabs Yellowbrick Python Scikit Learn Environment
- Environment:OpenHands OpenHands Third Party Runtime Credentials
- Environment:Openai CLIP PyTorch CUDA Runtime
- Environment:DataTalksClub Data engineering zoomcamp Kafka Confluent Environment
- Environment:Protectai Llm guard Python Runtime Dependencies
- Environment:Heibaiying BigData Notes Spark 2 4 Environment
- Environment:Snorkel team Snorkel PySpark
- Environment:Onnx Onnx Python Runtime Environment