Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent with the Leeroopedia MCP setup guide. Let it search docs, build plans, verify code, and diagnose failures on your behalf.
Go end-to-end. Leeroopedia gives your agent the knowledge. Kapso gives it the ability to act on it: research, experiment, and deploy.
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
__NOCACHE__
Workflows
- Workflow:Explodinggradients Ragas LLM Benchmarking
- Workflow:Pytorch Serve HuggingFace Transformer Serving
- Workflow:Google research Deduplicate text datasets Cross dataset deduplication
- Workflow:AUTOMATIC1111 Stable diffusion webui LoRA network application
- Workflow:DataExpert io Data engineer handbook PySpark Iceberg Job Execution
- Workflow:Wandb Weave Evaluation Pipeline
- Workflow:Apache Dolphinscheduler Datasource Connection Management
- Workflow:Huggingface Trl Direct Preference Optimization
- Workflow:Iterative Dvc Experiment Tracking
- Workflow:Togethercomputer Together python Image Generation
Principles
- Principle:Puppeteer Puppeteer HTTP Response Handling
- Principle:Pytorch Serve GRPC Communication
- Principle:Run llama Llama index Global Settings Configuration
- Principle:Langchain ai Langchain Public API Surface
- Principle:Eventual Inc Daft Temp Table Registration
- Principle:Openclaw Openclaw Docker Image Build
- Principle:Run llama Llama index Tool Definition
- Principle:Cleanlab Cleanlab Latent Noise Estimation
- Principle:Shiyu coder Kronos Tokenizer Loading
- Principle:Nautechsystems Nautilus trader Instrument Specification
Implementations
- Implementation:OWASP Www project top 10 for large language model applications VulnerabilityEntry Parse Sections
- Implementation:Risingwavelabs Risingwave Grafana Dashboard Generation
- Implementation:Huggingface Trl GRPOTrainer Train Loop
- Implementation:Getgauge Taiko Repl Initialize
- Implementation:OpenRLHF OpenRLHF Actor init
- Implementation:Apache Paimon RoaringBitmap
- Implementation:ArroyoSystems Arroyo Run Pipeline
- Implementation:Speechbrain Speechbrain Hparams Switchboard Transformer
- Implementation:AnswerDotAI RAGatouille RAGPretrainedModel From Pretrained
- Implementation:Apache Hudi Stop Demo Script
Heuristics
- Heuristic:Elevenlabs Elevenlabs python Audio Buffer Sizes
- Heuristic:DevExpress Testcafe Video Encoding Defaults
- Heuristic:Promptfoo Promptfoo Adaptive Concurrency Tuning
- Heuristic:LMCache LMCache Prefix Based Retrieval Pattern
- Heuristic:Eric mitchell Direct preference optimization FSDP Batch Size Per GPU
- Heuristic:FMInference FlexLLMGen Weight Compression 4bit
- Heuristic:Duckdb Duckdb Sanitizer Configuration
- Heuristic:Iamhankai Forest of Thought Self Correction Confidence Threshold
- Heuristic:Tensorflow Tfjs WASM Cross Origin Isolation
- Heuristic:Mlfoundations Open flamingo FSDP Manual Wrapping For Mixed Parameters
Environments
- Environment:InternLM Lmdeploy Build From Source
- Environment:Apache Druid Integration Test Docker
- Environment:DataTalksClub Data engineering zoomcamp PySpark Batch Environment
- Environment:OpenRLHF OpenRLHF CUDA GPU Environment
- Environment:Deepset ai Haystack GPU Device Environment
- Environment:Gretelai Gretel synthetics PyTorch CUDA Environment
- Environment:Apache Shardingsphere Etcd Cluster Coordination
- Environment:Ollama Ollama GPU Runtime
- Environment:Rapidsai Cuml Dask Distributed
- Environment:Microsoft Autogen LLM Provider API Keys