Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Sdv dev SDV Data quality evaluation
- Workflow:Huggingface Open r1 Dataset Pass Rate Filtering
- Workflow:Confident ai Deepeval LLM Tracing and Observability
- Workflow:Tensorflow Serving Batched Inference Pipeline
- Workflow:Treeverse LakeFS Write Audit Publish With Hooks
- Workflow:Risingwavelabs Risingwave CDC Data Replication
- Workflow:FMInference FlexLLMGen Text Completion API
- Workflow:CARLA simulator Carla Building from Source
- Workflow:Mistralai Client python Finetuning Job Management
- Workflow:Kserve Kserve Canary Rollout Deployment
Principles
- Principle:Apache Shardingsphere YAML Configuration Definition
- Principle:Rapidsai Cuml Synthetic Dataset Generation
- Principle:Spcl Graph of thoughts GoT Graph Topology Design
- Principle:FMInference FlexLLMGen Optimized Inference Engine
- Principle:Apache Druid Visual Query Execution
- Principle:Huggingface Datatrove JSONL Data Writing
- Principle:Vibrantlabsai Ragas LLM Configuration
- Principle:Apache Airflow RC Build Signing
- Principle:Deepspeedai DeepSpeed Inference Engine Init
- Principle:Sktime Pytorch forecasting DeepAR Model Instantiation
Implementations
- Implementation:Online ml River Drift NoDrift
- Implementation:Speechbrain Speechbrain Prepare CommonVoice SSL
- Implementation:Apache Kafka Kafka Server Start Script
- Implementation:Norrrrrrr lyn WAInjectBench text ensemble detect
- Implementation:BerriAI Litellm Lowest Cost Strategy
- Implementation:Online ml River Reco BiasedMF
- Implementation:Run llama Llama index AzureOpenAIFinetuneEngine
- Implementation:Cohere ai Cohere python ToolCallV2 Model
- Implementation:Huggingface Datasets Pdf
- Implementation:Datajuicer Data juicer ImageSubplotFilter
Heuristics
- Heuristic:Microsoft Onnxruntime Memory Recomputation Optimization
- Heuristic:Deepset ai Haystack BM25 Score Scaling
- Heuristic:Dotnet Machinelearning FastTree Default Hyperparameters
- Heuristic:Mage ai Mage ai Sorted Data Bookmark Strategy
- Heuristic:Isaac sim IsaacGymEnvs JIT Profiling Optimization
- Heuristic:Onnx Onnx Opset Version Selection
- Heuristic:Apache Shardingsphere Shadow Routing Hint First Fallback
- Heuristic:Spcl Graph of thoughts Budget Gated Benchmark Execution
- Heuristic:Duckdb Duckdb PR Submission Strategy
- Heuristic:Openai Openai node Stream Usage Interruption
Environments
- Environment:Marker Inc Korea AutoRAG Vector Database Backends
- Environment:Kubeflow Kubeflow Python KFP SDK Environment
- Environment:Openai Openai python Azure OpenAI
- Environment:Langfuse Langfuse PostgreSQL 17
- Environment:MarketSquare Robotframework browser CI GitHub Actions
- Environment:Testtimescaling Testtimescaling github io Python 3 Runtime
- Environment:DataExpert io Data engineer handbook PostgreSQL Docker Environment
- Environment:HKUDS AI Trader Python LangChain Runtime
- Environment:Allenai Open instruct CUDA GPU Training
- Environment:Kserve Kserve Leader Worker Set