Main Page
Welcome to Leeroopedia
Your ML & Data Knowledge Wiki. Best practices and expert-level knowledge for Machine Learning and Data Engineering, covering 1000+ frameworks and libraries from training to deployment.
Browse implementation patterns, configuration guides, debugging heuristics, and battle-tested defaults for frameworks like vLLM, DeepSpeed, Megatron-LM, FlashAttention, Triton, Unsloth, LangChain, and many more. Every page is structured so both humans and AI agents can find what they need fast.
Connect your AI coding agent. Plug Leeroopedia into your favorite coding agent, and let it build robust AI/ML systems autonomously:
- SuperML plugin — converts your AI coding agent into an expert ML engineer with agentic memory
- Leeroopedia MCP — search over best-practices and skills of ML/AI
- Kapso — experimentation platform for autonomous AI/ML software building
Browse by Category
| Category | Description | Browse |
|---|---|---|
| Workflows | Step-by-step processes and procedures | Browse All |
| Principles | Core ideas and foundational knowledge | Browse All |
| Implementations | Code-level details and modules | Browse All |
| Heuristics | Best practices and guidelines | Browse All |
| Environments | Setup and configuration guides | Browse All |
Explore Pages
Workflows
- Workflow:PacktPublishing LLM Engineers Handbook Model Evaluation
- Workflow:Huggingface Trl Reward Model Training
- Workflow:Cohere ai Cohere python AWS Bedrock Deployment
- Workflow:NVIDIA NeMo Aligner RLHF PPO Training
- Workflow:Trailofbits Fickling PyTorch Format Identification
- Workflow:FlagOpen FlagEmbedding Benchmark Evaluation
- Workflow:Triton inference server Server Custom Container Build
- Workflow:SqueezeAILab ETS ETS Experiment Pipeline
- Workflow:Microsoft Autogen Swarm Agent Handoff
- Workflow:Kubeflow Kubeflow Contributing To Kubeflow
Principles
- Principle:Datahub project Datahub Sample Data Ingestion
- Principle:Guardrails ai Guardrails Stream Chunk Processing
- Principle:CarperAI Trlx Synthetic Environment Design
- Principle:Romsto Speculative Decoding Input Tokenization
- Principle:Huggingface Datasets Dataset Sorting
- Principle:Speechbrain Speechbrain Speaker Embedding Precomputation
- Principle:Princeton nlp SimPO Response Post Processing
- Principle:Google deepmind Mujoco Passive Forces
- Principle:Ggml org Llama cpp LoRA Model Merging
- Principle:OpenHands OpenHands Post Migration Verification
Implementations
- Implementation:Infiniflow Ragflow UseFileRequest Hooks
- Implementation:Scikit learn Scikit learn FetchLfw
- Implementation:Online ml River Preprocessing RandomProjection
- Implementation:Openai Openai node Responses Resource
- Implementation:Ggml org Llama cpp Jinja Parser
- Implementation:Ray project Ray Ray Init For Actors
- Implementation:SeleniumHQ Selenium Closure Structs Map
- Implementation:Scikit learn Scikit learn Train Test Split
- Implementation:ArroyoSystems Arroyo Datafusion Types
- Implementation:Tensorflow Tfjs Pooling Layers
Heuristics
- Heuristic:Microsoft LoRA Selective LoRA QV Only
- Heuristic:LLMBook zh LLMBook zh github io Greedy Decoding Temperature Zero
- Heuristic:Risingwavelabs Risingwave Docker Memory Allocation
- Heuristic:Run llama Llama index Embedding Batch Size Tuning
- Heuristic:Liu00222 Open Prompt Injection Cosine Similarity Segmentation Threshold
- Heuristic:Infiniflow Ragflow Reranking Weight Tuning
- Heuristic:Apache Hudi Data Skipping Limitations
- Heuristic:Haifengl Smile Platform Native Library Selection
- Heuristic:Vibrantlabsai Ragas Analytics Silent Failure Pattern
- Heuristic:Lucidrains X transformers Numerical Stability Techniques
Environments
- Environment:Treeverse LakeFS Web UI Environment
- Environment:Bentoml BentoML Python Runtime
- Environment:Apache Dolphinscheduler Database Backend
- Environment:Explodinggradients Ragas LLM Provider Environment
- Environment:Heibaiying BigData Notes Kafka 2 2 Environment
- Environment:Vllm project Vllm CUDA
- Environment:Eric mitchell Direct preference optimization HuggingFace Transformers
- Environment:Dotnet Machinelearning OneDal Acceleration
- Environment:Google research Deduplicate text datasets Python HuggingFace Environment
- Environment:Haosulab ManiSkill Python SAPIEN Core