Implementation:Ggml org Ggml Cpu backend interface
Metadata
| Field | Value |
|---|---|
| Page Type | Implementation (Backend Interface) |
| Knowledge Sources | GGML |
| Domains | ML_Infrastructure, Tensor_Computing, CPU_Backend |
| Last Updated | 2025-05-15 12:00 GMT |
Overview
Implements the GGML backend interface for the CPU, providing device registration, buffer management, graph planning, and compute dispatch through the backend API.
Description
ggml-cpu.cpp is the bridge between GGML's abstract backend system and the CPU-specific compute implementation. It enables the CPU to be used interchangeably with GPU backends through a unified interface. Key components include:
- Backend context: Defines
ggml_backend_cpu_contextholding thread count, threadpool, work buffer, abort callback, anduse_refflag. - Backend interface (
ggml_backend_i): Implementsget_name(returns "CPU"),graph_plan_create/graph_plan_compute(wrappingggml_graph_plan/ggml_graph_compute), andgraph_computefor planless execution. - Extra buffer types: Collects accelerator buffer types (AMX, KleidiAI, SpaceMIT, repack) via
ggml_backend_cpu_get_extra_buffer_types(), enabling specialized tensor memory layouts for optimized kernels. - Device interface: Reports device properties (name, description, memory size via sysconf/sysctl), handles op support queries, and provides device-level buffer allocation.
- Backend registration: Exports
ggml_backend_cpu_reg()for the backend registry, andggml_backend_cpu_init()for direct instantiation.
Usage
Use ggml_backend_cpu_init() to create a CPU backend instance, or let the backend registry discover it automatically. The CPU backend is typically used as a fallback or primary backend for inference.
Code Reference
Source Location
GGML repo, file: src/ggml-cpu/ggml-cpu.cpp (701 lines).
Signature
// Create a CPU backend instance
ggml_backend_t ggml_backend_cpu_init(void);
// Get the CPU backend registry entry
ggml_backend_reg_t ggml_backend_cpu_reg(void);
// Get extra buffer types for accelerators (AMX, KleidiAI, repack, etc.)
std::vector<ggml_backend_buffer_type_t> & ggml_backend_cpu_get_extra_buffer_types();
Import
#include "ggml-cpu.h"
#include "ggml-backend.h"
I/O Contract
Inputs
| Parameter | Type | Required | Description |
|---|---|---|---|
| (none for init) | ggml_backend_cpu_init takes no parameters; default configuration is applied.
| ||
n_threads |
int |
Via setter | Set with ggml_backend_cpu_set_n_threads() after init.
|
threadpool |
ggml_threadpool_t |
Via setter | Set with ggml_backend_cpu_set_threadpool() after init.
|
abort_callback |
ggml_abort_callback |
Via setter | Set with ggml_backend_cpu_set_abort_callback() after init.
|
Outputs
| Output | Type | Description |
|---|---|---|
| Backend handle | ggml_backend_t |
Opaque handle to the CPU backend, usable with all ggml_backend_* functions.
|
| Registry entry | ggml_backend_reg_t |
Static singleton for backend auto-discovery. |
Usage Examples
Creating and Using a CPU Backend
#include "ggml-cpu.h"
#include "ggml-backend.h"
// Create CPU backend
ggml_backend_t cpu = ggml_backend_cpu_init();
// Configure threads
ggml_backend_cpu_set_n_threads(cpu, 8);
// Allocate a buffer for tensors
ggml_backend_buffer_t buf = ggml_backend_alloc_buffer(cpu, 64 * 1024 * 1024);
// ... build graph, allocate tensors, compute ...
ggml_backend_graph_compute(cpu, graph);
// Cleanup
ggml_backend_buffer_free(buf);
ggml_backend_free(cpu);
Related Pages
- Implementation:Ggml_org_Ggml_Cpu_compute_engine -- The C-level graph compute engine wrapped by this interface.
- Implementation:Ggml_org_Ggml_Cpu_amx_mmq -- AMX accelerator registered as an extra buffer type.
- Implementation:Ggml_org_Ggml_Cpu_kleidiai_backend -- KleidiAI accelerator registered as an extra buffer type.
- Implementation:Ggml_org_Ggml_Cpu_weight_repack -- Weight repacking registered as an extra buffer type.