Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Ggml org Ggml Cpu backend interface

From Leeroopedia


Metadata

Field Value
Page Type Implementation (Backend Interface)
Knowledge Sources GGML
Domains ML_Infrastructure, Tensor_Computing, CPU_Backend
Last Updated 2025-05-15 12:00 GMT

Overview

Implements the GGML backend interface for the CPU, providing device registration, buffer management, graph planning, and compute dispatch through the backend API.

Description

ggml-cpu.cpp is the bridge between GGML's abstract backend system and the CPU-specific compute implementation. It enables the CPU to be used interchangeably with GPU backends through a unified interface. Key components include:

  1. Backend context: Defines ggml_backend_cpu_context holding thread count, threadpool, work buffer, abort callback, and use_ref flag.
  2. Backend interface (ggml_backend_i): Implements get_name (returns "CPU"), graph_plan_create/graph_plan_compute (wrapping ggml_graph_plan/ggml_graph_compute), and graph_compute for planless execution.
  3. Extra buffer types: Collects accelerator buffer types (AMX, KleidiAI, SpaceMIT, repack) via ggml_backend_cpu_get_extra_buffer_types(), enabling specialized tensor memory layouts for optimized kernels.
  4. Device interface: Reports device properties (name, description, memory size via sysconf/sysctl), handles op support queries, and provides device-level buffer allocation.
  5. Backend registration: Exports ggml_backend_cpu_reg() for the backend registry, and ggml_backend_cpu_init() for direct instantiation.

Usage

Use ggml_backend_cpu_init() to create a CPU backend instance, or let the backend registry discover it automatically. The CPU backend is typically used as a fallback or primary backend for inference.

Code Reference

Source Location

GGML repo, file: src/ggml-cpu/ggml-cpu.cpp (701 lines).

Signature

// Create a CPU backend instance
ggml_backend_t ggml_backend_cpu_init(void);

// Get the CPU backend registry entry
ggml_backend_reg_t ggml_backend_cpu_reg(void);

// Get extra buffer types for accelerators (AMX, KleidiAI, repack, etc.)
std::vector<ggml_backend_buffer_type_t> & ggml_backend_cpu_get_extra_buffer_types();

Import

#include "ggml-cpu.h"
#include "ggml-backend.h"

I/O Contract

Inputs

Parameter Type Required Description
(none for init) ggml_backend_cpu_init takes no parameters; default configuration is applied.
n_threads int Via setter Set with ggml_backend_cpu_set_n_threads() after init.
threadpool ggml_threadpool_t Via setter Set with ggml_backend_cpu_set_threadpool() after init.
abort_callback ggml_abort_callback Via setter Set with ggml_backend_cpu_set_abort_callback() after init.

Outputs

Output Type Description
Backend handle ggml_backend_t Opaque handle to the CPU backend, usable with all ggml_backend_* functions.
Registry entry ggml_backend_reg_t Static singleton for backend auto-discovery.

Usage Examples

Creating and Using a CPU Backend

#include "ggml-cpu.h"
#include "ggml-backend.h"

// Create CPU backend
ggml_backend_t cpu = ggml_backend_cpu_init();

// Configure threads
ggml_backend_cpu_set_n_threads(cpu, 8);

// Allocate a buffer for tensors
ggml_backend_buffer_t buf = ggml_backend_alloc_buffer(cpu, 64 * 1024 * 1024);

// ... build graph, allocate tensors, compute ...
ggml_backend_graph_compute(cpu, graph);

// Cleanup
ggml_backend_buffer_free(buf);
ggml_backend_free(cpu);

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment