Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:ChenghaoMou Text dedup Benchmark Evaluation
- Workflow:Princeton nlp SimPO Model Inference
- Workflow:ArroyoSystems Arroyo Connection Setup
- Workflow:Datajuicer Data juicer Distributed Ray Processing
- Workflow:Lakeraai Pint benchmark Custom Dataset Benchmarking
- Workflow:Zai org CogVideo SAT Finetuning
- Workflow:Microsoft Onnxruntime On Device Training
- Workflow:Protectai Modelscan CLI Model Scanning
- Workflow:Explodinggradients Ragas LLM Benchmarking
- Workflow:Arize ai Phoenix Dataset and Experiment Lifecycle
Principles
- Principle:FMInference FlexLLMGen NVMe Disk Setup
- Principle:Nightwatchjs Nightwatch Page Object Usage In Tests
- Principle:ARISE Initiative Robosuite Joint Velocity Control
- Principle:Togethercomputer Together python Result Integration
- Principle:Langchain ai Langgraph Checkpointer Initialization
- Principle:Vespa engine Vespa RPM Package Creation
- Principle:Datahub project Datahub Emitter Instantiation
- Principle:Apache Druid Dimension Measure Configuration
- Principle:Ggml org Ggml Computation Graph Construction
- Principle:FlowiseAI Flowise Vector Store Query
Implementations
- Implementation:Mlc ai Web llm Tool Choice Request
- Implementation:CrewAIInc CrewAI Tavily Extractor Tool
- Implementation:SeleniumHQ Selenium Civetweb Core
- Implementation:Togethercomputer Together python Result Integration Pattern
- Implementation:Google deepmind Mujoco User Mesh
- Implementation:Googleapis Python genai Models Generate Images
- Implementation:Mlfoundations Open flamingo Create model and transforms
- Implementation:Scikit learn Scikit learn PlottingBase
- Implementation:PeterL1n BackgroundMattingV2 Load pretrained deeplabv3 state dict
- Implementation:Dagster io Dagster MonthlyPartitionsDefinition API
Heuristics
- Heuristic:Mlc ai Mlc llm GPU Memory Budget Tuning
- Heuristic:Google deepmind Mujoco MJX Benchmarking Tips
- Heuristic:Vllm project Vllm Batch Size Hardware Scaling
- Heuristic:Google research Deduplicate text datasets Ulimit File Descriptors For Merge
- Heuristic:ContextualAI HALOs FSDP Sampling Workaround
- Heuristic:DataTalksClub Data engineering zoomcamp Kafka Consumer Poll Timeout
- Heuristic:Onnx Onnx Shape Inference Limitations
- Heuristic:InternLM Lmdeploy KV Quantization Tradeoffs
- Heuristic:SeldonIO Seldon core Over Commit Memory Tip
- Heuristic:Lance format Lance Vector Index Tuning
Environments
- Environment:Bitsandbytes foundation Bitsandbytes ROCm AMD Environment
- Environment:Microsoft LoRA NLG Eval External Tools
- Environment:Langchain ai Langchain LangSmith Tracing Config
- Environment:Lance format Lance Rust Toolchain
- Environment:Open compass VLMEvalKit GPU CUDA Environment
- Environment:Duckdb Duckdb Extension Distribution Env
- Environment:Apache Dolphinscheduler Node Pnpm Runtime
- Environment:Speechbrain Speechbrain PyTorch CUDA Runtime
- Environment:Datahub project Datahub Frontend Build
- Environment:Google deepmind Mujoco MJX Warp CUDA Environment