Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Infiniflow Ragflow Knowledge Base Document Ingestion
- Workflow:Duckdb Duckdb Benchmark Execution
- Workflow:NVIDIA NeMo Aligner DPO Training
- Workflow:Online ml River Streaming Anomaly Detection
- Workflow:Huggingface Open r1 Dataset Pass Rate Filtering
- Workflow:Sgl project Sglang Offline Batch Inference
- Workflow:Obss Sahi COCO Dataset Slicing
- Workflow:Apache Kafka Docker Image Release
- Workflow:Apache Kafka Coordinator Runtime Lifecycle
- Workflow:FlagOpen FlagEmbedding Reranker Finetuning
Principles
- Principle:Ollama Ollama Model Alias Resolution
- Principle:Eventual Inc Daft Iceberg Reading
- Principle:Pyro ppl Pyro Discrete Posterior Decoding
- Principle:Mit han lab Llm awq Fused Attention Optimization
- Principle:PacktPublishing LLM Engineers Handbook Quantized Model Loading
- Principle:CARLA simulator Carla Sensor Data Streaming
- Principle:Fede1024 Rust rdkafka Async Message Production
- Principle:FlowiseAI Flowise Tool Management
- Principle:Heibaiying BigData Notes Hive Partitioning and Bucketing
- Principle:Roboflow Rf detr Detection Visualization
Implementations
- Implementation:Elevenlabs Elevenlabs python GetKnowledgeBaseSummaryUrlResponseModel
- Implementation:Ucbepic Docetl Pipeline Optimize
- Implementation:Ucbepic Docetl OutputPanel
- Implementation:Ggml org Llama cpp Json Partial
- Implementation:Apache Shardingsphere SchemaMetaDataPersistService Persist
- Implementation:NVIDIA NeMo Curator Metrics Utils
- Implementation:Speechbrain Speechbrain Train CommonVoice Transformer
- Implementation:NVIDIA TransformerEngine Custom GEMM
- Implementation:Protectai Llm guard Scan output
- Implementation:Vespa engine Vespa Application ABI Spec
Heuristics
- Heuristic:Openai Evals Event Batching Configuration
- Heuristic:Intel Ipex llm DeepSpeed Tensor Parallel Tips
- Heuristic:Iterative Dvc Path Performance Optimization
- Heuristic:PeterL1n BackgroundMattingV2 Mixed Precision Training
- Heuristic:ARISE Initiative Robomimic Video Recording Optimization
- Heuristic:PeterL1n BackgroundMattingV2 Training Batch Size And Resolution
- Heuristic:Scikit learn Scikit learn Warning Deprecated PassiveAggressive
- Heuristic:SeldonIO Seldon core Kafka Partition Throughput Tip
- Heuristic:Bigscience workshop Petals Short Inference Pool Merging
- Heuristic:Intel Ipex llm Use Cache Training Vs Inference
Environments
- Environment:Mlc ai Mlc llm WebGPU Browser Environment
- Environment:Google research Deduplicate text datasets Python TFDS Environment
- Environment:Apache Spark Release Build Environment
- Environment:Neuml Txtai Optional Extras Dependencies
- Environment:ThreeSR Awesome Inference Time Scaling Python Runtime Environment
- Environment:Eric mitchell Direct preference optimization Python Dependencies
- Environment:Kubeflow Pipelines Python SDK
- Environment:Heibaiying BigData Notes Kafka 2 2 Environment
- Environment:Google research Deduplicate text datasets Python HuggingFace Environment
- Environment:Datajuicer Data juicer GPU CUDA Environment