Implementation:Hiyouga LLaMA Factory Glaive Toolcall En Demo Data
| Knowledge Sources | |
|---|---|
| Domains | NLP, Training_Data |
| Last Updated | 2026-02-06 19:00 GMT |
Overview
glaive_toolcall_en_demo.json provides English tool/function calling demo data in ShareGPT format with tool definitions and multi-turn conversations for training models to invoke external functions.
Description
The file contains a JSON array of records, each with a conversations list and a tools string. The conversations follow ShareGPT format using role tags "from" with values "human", "gpt", "function_call", and "observation". The "function_call" turns contain JSON-serialized tool invocations with function name and arguments, while "observation" turns contain the JSON-serialized tool response. The tools field is a JSON string defining available function schemas following the OpenAI function calling format.
Some conversations include tool calls while others are purely conversational, demonstrating that the model should learn when to invoke tools and when to respond directly.
Usage
This demo dataset trains and tests function-calling capabilities in LLaMA Factory. Users reference it by name (glaive_toolcall_en_demo) with --stage sft to fine-tune models for tool use. The tool definitions in the tools field are injected as part of the system context during training.
Code Reference
Source Location
- Repository: Hiyouga_LLaMA_Factory
- File: data/glaive_toolcall_en_demo.json
Data Format
[
{
"conversations": [
{
"from": "human",
"value": "Hi, I have some ingredients and I want to cook something. Can you help me find a recipe?"
},
{
"from": "gpt",
"value": "Of course! I can help you with that. Please tell me what ingredients you have."
},
{
"from": "human",
"value": "I have chicken, bell peppers, and rice."
},
{
"from": "function_call",
"value": "{\"name\": \"search_recipes\", \"arguments\": {\"ingredients\": [\"chicken\", \"bell peppers\", \"rice\"]}}"
},
{
"from": "observation",
"value": "{\"recipes\": [{\"name\": \"Chicken and Bell Pepper Stir Fry\", ...}]}"
},
{
"from": "gpt",
"value": "I found two recipes for you..."
}
],
"tools": "[{\"name\": \"search_recipes\", \"description\": \"Search for recipes based on ingredients\", \"parameters\": {\"type\": \"object\", \"properties\": {\"ingredients\": {\"type\": \"array\", \"items\": {\"type\": \"string\"}, \"description\": \"The ingredients to search for\"}}, \"required\": [\"ingredients\"]}}]"
}
]
I/O Contract
Schema
| Field | Type | Required | Description |
|---|---|---|---|
| conversations | array | Yes | List of conversation turns with "from" and "value" fields
|
| tools | string | Yes | JSON-serialized array of tool/function definitions following OpenAI function schema |
Conversation Roles
| Role (from) | Description |
|---|---|
| human | User message |
| gpt | Assistant response (text only, no tool call) |
| function_call | Tool invocation with JSON payload containing name and arguments
|
| observation | Tool execution result as JSON |
Dataset Registry Entry
| Property | Value |
|---|---|
| Key | glaive_toolcall_en_demo
|
| file_name | glaive_toolcall_en_demo.json
|
| formatting | sharegpt |
| columns.messages | conversations |
| columns.tools | tools |
| Lines | 9158 |
Usage Examples
# Train a model with function-calling capabilities
# llamafactory-cli train \
# --dataset glaive_toolcall_en_demo \
# --stage sft \
# --model_name_or_path meta-llama/Llama-2-7b-hf \
# --output_dir output/toolcall_demo
# Loading and inspecting the data
import json
with open("data/glaive_toolcall_en_demo.json", "r", encoding="utf-8") as f:
data = json.load(f)
print(f"Number of conversations: {len(data)}")
sample = data[0]
print(f"Turns in first conversation: {len(sample['conversations'])}")
tools = json.loads(sample['tools'])
print(f"Number of tools defined: {len(tools)}")
for tool in tools:
print(f" Tool: {tool['name']} - {tool['description']}")
Related Pages
- Hiyouga_LLaMA_Factory_Glaive_Toolcall_Zh_Demo_Data - Chinese version of the tool-calling demo dataset
- Hiyouga_LLaMA_Factory_Dataset_Info_Registry - Central dataset registry that indexes this file
- Hiyouga_LLaMA_Factory_Alpaca_En_Demo_Data - English SFT demo data (non-tool-calling)