Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:Microsoft LoRA LoRA Integration
- Workflow:Ggml org Ggml Backend Accelerated Computation
- Workflow:MarketSquare Robotframework browser Installation and Setup
- Workflow:Apache Druid Batch Data Ingestion
- Workflow:Mistralai Client python Chat Completion
- Workflow:FlagOpen FlagEmbedding Benchmark Evaluation
- Workflow:Arize ai Phoenix Prompt Management Pipeline
- Workflow:Infiniflow Ragflow Knowledge Base Document Ingestion
- Workflow:Apache Flink Stream File Compaction
- Workflow:Arize ai Phoenix Span Annotation Pipeline
Principles
- Principle:Ggml org Ggml CPU Feature Detection
- Principle:TobikoData Sqlmesh PR Environment Creation
- Principle:Apache Shardingsphere SPI Refresher Loading
- Principle:EvolvingLMMs Lab Lmms eval Model Configuration
- Principle:Apache Flink Enriched Row Composition
- Principle:FMInference FlexLLMGen Resource Cleanup
- Principle:NVIDIA NeMo Curator Text Cleaning and Normalization
- Principle:OWASP Www project top 10 for large language model applications Supplemental Content Translation
- Principle:Langfuse Langfuse OTel Input Output Extraction
- Principle:Scikit learn contrib Imbalanced learn Value Difference Metric
Implementations
- Implementation:Kubeflow Pipelines XGBoost Train Incremental
- Implementation:Volcengine Verl Multimodal Rollout Request
- Implementation:Avhz RustQuant FFT
- Implementation:SeleniumHQ Selenium FirefoxCommandContext
- Implementation:Iterative Dvc Repo Get
- Implementation:Open compass VLMEvalKit MMHelix Aquarium Eval
- Implementation:Ggml org Ggml Ggml opt fit
- Implementation:Alibaba ROLL DPOPipeline Val
- Implementation:Microsoft LoRA Legacy Run SWAG
- Implementation:OpenHands OpenHands OrgService Get Owner Role
Heuristics
- Heuristic:Apache Airflow DAG Complexity Reduction
- Heuristic:Apache Flink False Positive Availability Optimization
- Heuristic:NVIDIA DALI Warning Deprecated C API V1 Functions
- Heuristic:Openai Openai python Fine Tuning Data Preparation Tips
- Heuristic:SeleniumHQ Selenium Warning Deprecated Proxy FTP Methods
- Heuristic:Google deepmind Dm control Physics Timestep Configuration
- Heuristic:OpenHands OpenHands Redis Distributed Locking
- Heuristic:Run llama Llama index Evaluator LLM Selection
- Heuristic:ClickHouse ClickHouse Jemalloc Production Requirement
- Heuristic:Lucidrains X transformers Gradient Clipping And Accumulation
Environments
- Environment:Ollama Ollama CGo Runtime
- Environment:Huggingface Alignment handbook Python Transformers
- Environment:Predibase Lorax Python Server Dependencies
- Environment:Snorkel team Snorkel Dask Distributed
- Environment:Princeton nlp Tree of thought llm Python OpenAI
- Environment:Eventual Inc Daft AI Provider Dependencies
- Environment:Vespa engine Vespa Cosign Sigstore Signing
- Environment:Haotian liu LLaVA Python CUDA Training Environment
- Environment:ThreeSR Awesome Inference Time Scaling GitHub Account Environment
- Environment:Dagster io Dagster GRPC Communication