Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Iterative Dvc Data Tracking
- Workflow:Fastai Fastbook NLP Text Classification
- Workflow:BerriAI Litellm Proxy Server Deployment
- Workflow:Huggingface Peft Seq2Seq AdaLoRA Finetuning
- Workflow:Intel Ipex llm RAG With LangChain
- Workflow:Gretelai Gretel synthetics ACTGAN Tabular Synthesis
- Workflow:Cohere ai Cohere python AWS Bedrock Deployment
- Workflow:Openclaw Openclaw Gateway Operations And Diagnostics
- Workflow:Langchain ai Langchain Streaming Responses
- Workflow:Huggingface Datasets Dataset Streaming
Principles
- Principle:Deepspeedai DeepSpeed Op Builder System
- Principle:TA Lib Ta lib python Streaming Indicator Computation
- Principle:DataExpert io Data engineer handbook Event Tracking
- Principle:Vllm project Vllm Server Metrics Monitoring
- Principle:Google research Deduplicate text datasets Cross Dataset Duplicate Detection
- Principle:Openai CLIP Image Feature Encoding
- Principle:Lm sys FastChat Seq2Seq SFT Training
- Principle:Openai Openai python Chat Response Processing
- Principle:Online ml River Estimator Base Architecture
- Principle:Astronomer Astronomer cosmos Operator Arguments Configuration
Implementations
- Implementation:Huggingface Diffusers DreamBooth Args
- Implementation:Webdriverio Webdriverio Appium Types
- Implementation:Scikit learn Scikit learn Perceptron
- Implementation:Evidentlyai Evidently Legacy Data Quality Preset
- Implementation:Langfuse Langfuse DatasetRunItemUpsertQueue
- Implementation:FlowiseAI Flowise DashboardMenu
- Implementation:EvolvingLMMs Lab Lmms eval OVOBench Utils
- Implementation:Openai Openai python Shared Response Format JSON Schema
- Implementation:Open compass VLMEvalKit LENS Utils
- Implementation:NVIDIA TransformerEngine Prepare TE Modules For FSDP
Heuristics
- Heuristic:Hpcaitech ColossalAI CUDA Device Max Connections Tip
- Heuristic:Mlflow Mlflow Model Signature Inference Tips
- Heuristic:Microsoft Agent framework Declaration Only Tools Pattern
- Heuristic:Fastai Fastbook Discriminative Learning Rates
- Heuristic:Explodinggradients Ragas Failed Metrics Return NaN
- Heuristic:Isaac sim IsaacGymEnvs Determinism Performance Tradeoff
- Heuristic:Mbzuai oryx Awesome LLM Post training Paper Deduplication Via Dict
- Heuristic:Datajuicer Data juicer Checkpoint Resumption Strategy
- Heuristic:Ggml org Ggml Memory Allocation Strategy
- Heuristic:Isaac sim IsaacGymEnvs DR Setup Only Flag
Environments
- Environment:Huggingface Trl Quantization Environment
- Environment:Axolotl ai cloud Axolotl Python Runtime
- Environment:Roboflow Rf detr Python GPU Environment
- Environment:Kubeflow Kubeflow Kubectl Kustomize CLI Environment
- Environment:Spotify Luigi Tornado Web Server
- Environment:Testtimescaling Testtimescaling github io GitHub Actions Runner
- Environment:Mit han lab Llm awq VILA Multimodal Environment
- Environment:Apache Kafka JVM Runtime Environment
- Environment:Huggingface Datatrove S3 Storage Environment
- Environment:FMInference FlexLLMGen CUDA GPU