Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:PrefectHQ Prefect Dbt Model Orchestration
- Workflow:Sgl project Sglang ModelOpt Quantization And Export
- Workflow:Apache Flink File Sink Pipeline
- Workflow:OpenRLHF OpenRLHF DPO Training
- Workflow:ArroyoSystems Arroyo Checkpoint Recovery
- Workflow:Microsoft BIPIA White Box Defense Finetuning
- Workflow:Tensorflow Tfjs GPT2 Text Generation
- Workflow:Microsoft Onnxruntime Distributed Model Training
- Workflow:Haosulab ManiSkill Motion Planning Demo Generation
- Workflow:Ucbepic Docetl Python API Pipeline
Principles
- Principle:OpenHands OpenHands Progress Notification
- Principle:Predibase Lorax Continuous Batching Inference
- Principle:DataExpert io Data engineer handbook Event Tracking
- Principle:Confident ai Deepeval Component Instrumentation
- Principle:Scikit learn contrib Imbalanced learn Combined Over Under Sampling Tomek
- Principle:Google deepmind Mujoco Rendering Resource Cleanup
- Principle:Promptfoo Promptfoo CLI Utilities
- Principle:Neuml Txtai Late Interaction Retrieval
- Principle:Sktime Pytorch forecasting Baseline Forecasting
- Principle:Run llama Llama index Finetuned Model Retrieval
Implementations
- Implementation:Openai Openai node Ecosystem CLI
- Implementation:FlagOpen FlagEmbedding Training Data JSONL Format
- Implementation:Isaac sim IsaacGymEnvs Launch Rlg Hydra Inference
- Implementation:Apache Druid CellFilterMenu
- Implementation:Hiyouga LLaMA Factory Alpaca Zh Demo Data
- Implementation:Hpcaitech ColossalAI Train ORPO Script
- Implementation:Deepspeedai DeepSpeed InferenceEngine Forward
- Implementation:BerriAI Litellm Responses API
- Implementation:Apache Paimon FormatAvroReader
- Implementation:Deepset ai Haystack SentenceTransformersTextEmbedder
Heuristics
- Heuristic:ChenghaoMou Text dedup Fingerprint Batch Size One
- Heuristic:Google deepmind Mujoco MJX Feature Compatibility
- Heuristic:Lm sys FastChat Flash Attention GPU Requirements
- Heuristic:Microsoft BIPIA LLAMA Pad Token Workaround
- Heuristic:Spotify Luigi Batch Parameter Aggregation
- Heuristic:Explodinggradients Ragas LLM Temperature Defaults
- Heuristic:Gretelai Gretel synthetics Memory Chunking For Normalization
- Heuristic:Langchain ai Langgraph Checkpointer Selection Guide
- Heuristic:Huggingface Datatrove VLLM Startup Optimization
- Heuristic:Rapidsai Cuml Quantile Split Differences
Environments
- Environment:Unslothai Unsloth Python Transformers
- Environment:Huggingface Alignment handbook Python TRL
- Environment:Facebookresearch Habitat lab CUDA GPU Training Environment
- Environment:Fede1024 Rust rdkafka Kafka Broker Runtime
- Environment:Getgauge Taiko Linux System Libraries
- Environment:Kserve Kserve Leader Worker Set
- Environment:Promptfoo Promptfoo Python Runtime
- Environment:Ggml org Ggml CUDA GPU Environment
- Environment:Online ml River Build Toolchain
- Environment:Huggingface Transformers Flash Attention 2 Env