Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Volcengine Verl GRPO Training Pipeline
- Workflow:NVIDIA DALI Image Preprocessing Pipeline
- Workflow:Triton inference server Server Custom Container Build
- Workflow:Huggingface Trl PPO RLHF Training
- Workflow:Explodinggradients Ragas RAG Evaluation
- Workflow:ChenghaoMou Text dedup Benchmark Evaluation
- Workflow:Openai Openai node Structured Output Parsing
- Workflow:Fede1024 Rust rdkafka Mock Cluster Testing
- Workflow:Kubeflow Kubeflow Release Management
- Workflow:Mage ai Mage ai Building a New Source Connector
Principles
- Principle:Truera Trulens Dashboard Visualization
- Principle:Apache Flink Split Based Record Reading
- Principle:CrewAIInc CrewAI Crew Integration In Flow
- Principle:Tencent Ncnn Neural Speech Synthesis
- Principle:ARISE Initiative Robosuite Joint Position Control
- Principle:Huggingface Diffusers Instance Data Collection
- Principle:CARLA simulator Carla Car Following Model
- Principle:Huggingface Datasets Pandas Dataset Building
- Principle:Pytorch Serve Model Registration
- Principle:Cleanlab Cleanlab Identifier Column Issue Detection
Implementations
- Implementation:Microsoft Playwright Page Network Events
- Implementation:Mlc ai Mlc llm MLCEngine Generate
- Implementation:Helicone Helicone LogController LogRequests
- Implementation:FlowiseAI Flowise CredentialsView
- Implementation:Ggml org Llama cpp Conversion Pip Dependencies
- Implementation:Snorkel team Snorkel SliceAwareClassifier Score Slices
- Implementation:Mit han lab Llm awq Split and repack
- Implementation:Avhz RustQuant Interpolator Trait
- Implementation:EvolvingLMMs Lab Lmms eval TUI Server
- Implementation:Ggml org Ggml Gpt2 graph
Heuristics
- Heuristic:LLMBook zh LLMBook zh github io LoRA Initialization Strategy
- Heuristic:FMInference FlexLLMGen Sequence Length Alignment
- Heuristic:Lance format Lance Vector Index Tuning
- Heuristic:Microsoft DeepSpeedExamples ZeRO Inference Throughput Tuning
- Heuristic:Snorkel team Snorkel DataParallel Default Behavior
- Heuristic:Apache Airflow Memory Management Tips
- Heuristic:Explodinggradients Ragas Embedding Batch Size Tuning
- Heuristic:FMInference FlexLLMGen OOM Memory Management
- Heuristic:Norrrrrrr lyn WAInjectBench L2 Normalize CLIP Embeddings
- Heuristic:Openai Openai python Fine Tuning Data Preparation Tips
Environments
- Environment:Ollama Ollama Go Runtime
- Environment:Farama Foundation Gymnasium Box2D Physics Backend
- Environment:Openai Openai python Azure OpenAI
- Environment:Protectai Modelscan H5py Optional
- Environment:Ollama Ollama GPU Runtime
- Environment:Mlfoundations Open flamingo PyTorch CUDA Distributed
- Environment:Tensorflow Serving Python Client Environment
- Environment:Apache Spark Kubernetes Runtime
- Environment:Vllm project Vllm AArch64 CPU
- Environment:DistrictDataLabs Yellowbrick Optional NLP Dependencies