Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:ArroyoSystems Arroyo Checkpoint Recovery
- Workflow:Nightwatchjs Nightwatch E2E Test Authoring
- Workflow:NVIDIA NeMo Curator Semantic Deduplication
- Workflow:Gretelai Gretel synthetics LSTM Text Generation
- Workflow:Deepseek ai Janus Rectified Flow Image Generation
- Workflow:Webdriverio Webdriverio Cucumber BDD Testing
- Workflow:Haotian liu LLaVA Web Demo Deployment
- Workflow:Helicone Helicone LLM Response Normalization
- Workflow:Apache Beam Twister2 Batch Execution
- Workflow:NVIDIA NeMo Curator Fuzzy Deduplication
Principles
- Principle:Apache Shardingsphere YAML Configuration Definition
- Principle:MaterializeInc Materialize Service Dependency Configuration
- Principle:Apache Paimon Lazy Blob Loading
- Principle:Ollama Ollama Architecture Detection
- Principle:Heibaiying BigData Notes Flink Kafka Source
- Principle:LaurentMazare Tch rs Backpropagation Training
- Principle:OpenGVLab InternVL Weighted Multi Dataset Training
- Principle:Unslothai Unsloth Project Configuration
- Principle:Infiniflow Ragflow Knowledge Base Creation
- Principle:Truera Trulens Method Instrumentation
Implementations
- Implementation:Helicone Helicone Unified Model Registry
- Implementation:Truera Trulens Block Input Decorator
- Implementation:Huggingface Datatrove DummyInferenceServer
- Implementation:Treeverse LakeFS Java SDK Model ImportCreation
- Implementation:Mlflow Mlflow Remove Experimental Decorators
- Implementation:InternLM Lmdeploy Impl Simt
- Implementation:LMCache LMCache Analyze Chunk Hashes
- Implementation:Langchain ai Langgraph Runtime Class
- Implementation:Turboderp org Exllamav2 ExLlamaV2PrefixFilter
- Implementation:Apache Airflow DagFileProcessorManager Discovery
Heuristics
- Heuristic:Google research Deduplicate text datasets Ulimit File Descriptors For Merge
- Heuristic:InternLM Lmdeploy Backend Selection Strategy
- Heuristic:Scikit learn Scikit learn Data Leakage Prevention
- Heuristic:Scikit learn contrib Imbalanced learn Sampling Strategy Selection
- Heuristic:Google research Deduplicate text datasets HACKSIZE Overlap Buffer
- Heuristic:Langchain ai Langchain Warning Deprecated Langchain Classic
- Heuristic:Langchain ai Langchain Error Context Preservation
- Heuristic:Alibaba ROLL KL Coefficient Tuning
- Heuristic:Onnx Onnx Shape Inference Limitations
- Heuristic:Protectai Llm guard ONNX Runtime Optimization
Environments
- Environment:Marker Inc Korea AutoRAG Python 3 10 Runtime
- Environment:Eric mitchell Direct preference optimization Python Dependencies
- Environment:BerriAI Litellm Python Runtime
- Environment:Microsoft LoRA NLU Conda Environment
- Environment:CARLA simulator Carla Python API Runtime
- Environment:Haifengl Smile Quarkus Serve Environment
- Environment:Huggingface Open r1 Slurm Cluster
- Environment:Vllm project Vllm CUDA GPU Runtime
- Environment:Datahub project Datahub Spark Lineage Environment
- Environment:Openai Openai python Voice Helpers