Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Ggml org Llama cpp Llama Model Load For Embeddings

From Leeroopedia
Revision as of 12:40, 16 February 2026 by Admin (talk | contribs) (Auto-imported from implementations/Ggml_org_Llama_cpp_Llama_Model_Load_For_Embeddings.md)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Field Value
Implementation Name Llama Model Load For Embeddings
Doc Type API Doc
Domain Model Loading, Embedding Configuration
Description Model loading with ctx_params.embeddings = true and pooling type configuration for embedding extraction
Related Workflow Embedding_Extraction

Overview

Description

The Llama Model Load For Embeddings implementation documents the API and configuration required to load a model in embedding extraction mode. This involves setting the embeddings flag in the context parameters, selecting the appropriate pooling type from the llama_pooling_type enum, and loading the model with llama_model_load_from_file(). The embedding example program (examples/embedding/embedding.cpp) demonstrates the canonical pattern for this configuration.

Usage

#include "llama.h"

// Enable embedding extraction in context parameters
llama_context_params ctx_params = llama_context_default_params();
ctx_params.embeddings = true;  // extract embeddings (together with logits)

// Load model
llama_model_params model_params = llama_model_default_params();
llama_model * model = llama_model_load_from_file("model.gguf", model_params);

// Create context with embedding support
llama_context * ctx = llama_init_from_model(model, ctx_params);

Code Reference

Field Value
Source Location (model load) include/llama.h:450-452
Source Location (embeddings field) include/llama.h:363
Source Location (pooling types) include/llama.h:168-175
Import #include "llama.h"

Model loading function signature:

// Load the model from a file
// If the file is split into multiple parts, the file name must follow this pattern: <name>-%05d-of-%05d.gguf
// If the split file name does not follow this pattern, use llama_model_load_from_splits
LLAMA_API struct llama_model * llama_model_load_from_file(
                         const char * path_model,
          struct llama_model_params   params);

Embeddings field in context parameters:

struct llama_context_params {
    // ...
    bool embeddings;  // if true, extract embeddings (together with logits)
    bool offload_kqv; // offload the KQV ops (including the KV cache) to GPU
    // ...
};

Pooling type enumeration:

enum llama_pooling_type {
    LLAMA_POOLING_TYPE_UNSPECIFIED = -1,
    LLAMA_POOLING_TYPE_NONE = 0,
    LLAMA_POOLING_TYPE_MEAN = 1,
    LLAMA_POOLING_TYPE_CLS  = 2,
    LLAMA_POOLING_TYPE_LAST = 3,
    LLAMA_POOLING_TYPE_RANK = 4, // used by reranking models to attach the classification head to the graph
};

Canonical embedding setup from examples/embedding/embedding.cpp:

int main(int argc, char ** argv) {
    common_params params;

    if (!common_params_parse(argc, argv, params, LLAMA_EXAMPLE_EMBEDDING)) {
        return 1;
    }

    common_init();
    params.embedding = true;

    // for non-causal models, batch size must be equal to ubatch size
    if (params.attention_type != LLAMA_ATTENTION_TYPE_CAUSAL) {
        params.n_ubatch = params.n_batch;
    }

    llama_backend_init();
    llama_numa_init(params.numa);

    // load the model
    auto llama_init = common_init_from_params(params);
    auto * model = llama_init->model();
    auto * ctx = llama_init->context();

    const enum llama_pooling_type pooling_type = llama_pooling_type(ctx);
    // ...
}

I/O Contract

Direction Description
Input Model file path (GGUF format), llama_model_params, llama_context_params with embeddings = true
Output Initialized llama_model * and llama_context * configured for embedding extraction
Preconditions Valid GGUF model file; sufficient memory for model weights; for embedding models, attention_type should not be CAUSAL
Error Handling Returns NULL if model file cannot be loaded; llama_pooling_type() returns the effective pooling type (may differ from requested if model metadata overrides)

Pooling type selection guide:

Pooling Type Enum Value Use Case Output Shape
None LLAMA_POOLING_TYPE_NONE (0) Per-token embeddings [n_tokens, n_embd]
Mean LLAMA_POOLING_TYPE_MEAN (1) Sentence embeddings (average) [n_sequences, n_embd]
CLS LLAMA_POOLING_TYPE_CLS (2) BERT-style first-token embedding [n_sequences, n_embd]
Last LLAMA_POOLING_TYPE_LAST (3) Causal model embeddings [n_sequences, n_embd]
Rank LLAMA_POOLING_TYPE_RANK (4) Reranking/classification scores [n_sequences, n_cls_out]

Usage Examples

CLI embedding extraction:

# Using the embedding example (--embedding flag set automatically)
./llama-embedding -m model.gguf -p "Hello world\nGoodbye world"

# Using the server with embedding mode
./llama-server -m embedding-model.gguf --embedding --port 8081

Programmatic loading with explicit pooling type:

common_params params;
params.embedding = true;
params.pooling_type = LLAMA_POOLING_TYPE_MEAN;  // explicit mean pooling
params.n_batch = 2048;
params.n_ubatch = 2048;  // must equal n_batch for non-causal models

auto llama_init = common_init_from_params(params);

Server-side embedding validation (from server.cpp):

// validate batch size for embeddings
if (params.embedding && params.n_batch > params.n_ubatch) {
    LOG_WRN("%s: setting n_batch = n_ubatch = %d to avoid assertion failure\n",
            __func__, params.n_ubatch);
    params.n_batch = params.n_ubatch;
}

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment