Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:ChenghaoMou Text dedup Bloom Filter Deduplication
- Workflow:Lance format Lance Vector Search Pipeline
- Workflow:Lance format Lance Full Text Search
- Workflow:Webdriverio Webdriverio Cucumber BDD Testing
- Workflow:DistrictDataLabs Yellowbrick Regression Model Evaluation
- Workflow:ThreeSR Awesome Inference Time Scaling Manual Paper Contribution
- Workflow:Huggingface Optimum Accelerated Inference Pipeline
- Workflow:Tensorflow Serving Model Version Management
- Workflow:Langfuse Langfuse Evaluation pipeline
- Workflow:Kubeflow Pipelines Iterative Model Training
Principles
- Principle:Allenai Open instruct GRPO Checkpointing
- Principle:Nightwatchjs Nightwatch Configuration Setup
- Principle:LLMBook zh LLMBook zh github io Pretraining Dataset Preparation
- Principle:Sgl project Sglang Distributed Environment Setup
- Principle:Avhz RustQuant Error Handling
- Principle:AUTOMATIC1111 Stable diffusion webui UI Accessibility
- Principle:Guardrails ai Guardrails Structured Output Consumption
- Principle:Openai Whisper Word Boundary Detection
- Principle:Getgauge Taiko Gauge Environment Configuration
- Principle:SeleniumHQ Selenium WebDriver Session For CDP
Implementations
- Implementation:Eventual Inc Daft Decode Image
- Implementation:Turboderp org Exllamav2 ExLlamaV2DynamicGenerator Iterate
- Implementation:Sktime Pytorch forecasting NormalDistributionLoss
- Implementation:Huggingface Diffusers ModelMixin From Pretrained Frozen
- Implementation:Obss Sahi Legacy NMS Postprocess
- Implementation:Arize ai Phoenix Legacy MistralAIModel
- Implementation:Alibaba ROLL DPOTrainer
- Implementation:FlowiseAI Flowise ChatPopUp
- Implementation:Neuml Txtai AudioStream
- Implementation:Dotnet Machinelearning Binary Classification Trainers
Heuristics
- Heuristic:Interpretml Interpret EBM Hyperparameter Tuning Guide
- Heuristic:Wandb Weave Batch Processing Tuning
- Heuristic:Datahub project Datahub Validation Across All APIs
- Heuristic:Unslothai Unsloth VLLM Memory Utilization
- Heuristic:Huggingface Peft Warning Deprecated Bone
- Heuristic:Lakeraai Pint benchmark Chunking Stride 25 Percent Overlap
- Heuristic:Mlfoundations Open flamingo Loss Masking Strategy
- Heuristic:Trailofbits Fickling Force Flag Bypass
- Heuristic:PrefectHQ Prefect Retry Backoff Strategy
- Heuristic:Openai Openai node Retry Backoff Configuration
Environments
- Environment:Google research Deduplicate text datasets Python TFDS Environment
- Environment:BerriAI Litellm Redis Cache Backend
- Environment:Huggingface Datasets Search Dependencies
- Environment:Risingwavelabs Risingwave Docker Deployment Environment
- Environment:Apache Airflow Kubernetes Helm Environment
- Environment:Fastai Fastbook CUDA GPU Environment
- Environment:NVIDIA DALI CMake Build Environment
- Environment:OpenBMB UltraFeedback vLLM Multi GPU Environment
- Environment:PrefectHQ Prefect AI Integration Credentials
- Environment:Mlc ai Mlc llm TVM Runtime Environment