Implementation:EvolvingLMMs Lab Lmms eval Crop Video MCP Server: Difference between revisions
Auto-imported from implementations/EvolvingLMMs_Lab_Lmms_eval_Crop_Video_MCP_Server.md |
Sync from local file |
||
| Line 1: | Line 1: | ||
{{PageInfo|type=Implementation|title=EvolvingLMMs_Lab_Lmms_eval_Crop_Video_MCP_Server}} | {{PageInfo|type=Implementation|title=EvolvingLMMs_Lab_Lmms_eval_Crop_Video_MCP_Server}} | ||
'''File''': <code>examples/mcp_server/crop_video_mcp_server.py</code> | |||
'''Principle''': [[MCP_Tool_Integration]] | |||
== Overview == | |||
The Crop Video MCP Server is a FastMCP-based server that provides video cropping functionality through the Model Context Protocol. It exposes a | The Crop Video MCP Server is a FastMCP-based server that provides video cropping functionality through the Model Context Protocol. It exposes a <code>crop_video</code> tool that extracts frames from a specified time range in a video file and returns them as a sequence of base64-encoded images. | ||
== Key Components == | |||
=== 1. Server Initialization === | |||
<syntaxhighlight lang="python"> | |||
app = FastMCP("Video Tools MCP Server", "0.1.0") | app = FastMCP("Video Tools MCP Server", "0.1.0") | ||
</syntaxhighlight> | |||
Creates a FastMCP application instance with name and version identifier. | Creates a FastMCP application instance with name and version identifier. | ||
=== 2. Logging Configuration === | |||
<syntaxhighlight lang="python"> | |||
logging.basicConfig( | logging.basicConfig( | ||
level=logging.INFO, | level=logging.INFO, | ||
| Line 22: | Line 22: | ||
handlers=[logging.FileHandler("/tmp/mcp_server_debug.log"), logging.StreamHandler()], | handlers=[logging.FileHandler("/tmp/mcp_server_debug.log"), logging.StreamHandler()], | ||
) | ) | ||
</syntaxhighlight> | |||
Sets up dual logging to both file and console for debugging server operations. | Sets up dual logging to both file and console for debugging server operations. | ||
=== 3. Tool Implementation === | |||
==== crop_video Function ==== | |||
<syntaxhighlight lang="python"> | |||
@app.tool(name="crop_video", description="Crop a video to a specified duration.") | @app.tool(name="crop_video", description="Crop a video to a specified duration.") | ||
def crop_video( | def crop_video( | ||
| Line 35: | Line 35: | ||
end_time: Annotated[float, Field(description="End time in seconds, must be > start_time")] = None, | end_time: Annotated[float, Field(description="End time in seconds, must be > start_time")] = None, | ||
) -> list[ImageContent]: | ) -> list[ImageContent]: | ||
</syntaxhighlight> | |||
'''Parameters''': | |||
- | - <code>video_path</code> (str): Path to the video file to process | ||
- | - <code>start_time</code> (float): Start time in seconds for cropping | ||
- | - <code>end_time</code> (float): End time in seconds (must be greater than start_time) | ||
'''Returns''': List of ImageContent objects containing extracted frames as base64 PNG images | |||
'''Key Operations''': | |||
1. | 1. '''Parameter Validation''': | ||
- Checks all required parameters are provided | - Checks all required parameters are provided | ||
- Validates parameter values (non-negative times, end > start) | - Validates parameter values (non-negative times, end > start) | ||
- Verifies video file existence | - Verifies video file existence | ||
2. | 2. '''Video Duration Verification''': | ||
```python | ```python | ||
cap = cv2.VideoCapture(video_path) | cap = cv2.VideoCapture(video_path) | ||
| Line 60: | Line 60: | ||
Uses OpenCV to check video duration and validate time range | Uses OpenCV to check video duration and validate time range | ||
3. | 3. '''Frame Extraction''': | ||
```python | ```python | ||
video_ele = { | video_ele = { | ||
| Line 76: | Line 76: | ||
Uses qwen_vl_utils.fetch_video to extract frames at 1 FPS | Uses qwen_vl_utils.fetch_video to extract frames at 1 FPS | ||
4. | 4. '''Image Encoding''': | ||
```python | ```python | ||
video_frames = video_frames.to(torch.uint8) | video_frames = video_frames.to(torch.uint8) | ||
| Line 91: | Line 91: | ||
Converts frames to PIL images, encodes as PNG, and wraps in ImageContent | Converts frames to PIL images, encodes as PNG, and wraps in ImageContent | ||
=== 4. Server Entry Point === | |||
<syntaxhighlight lang="python"> | |||
if __name__ == "__main__": | if __name__ == "__main__": | ||
app.run() | app.run() | ||
</syntaxhighlight> | |||
Launches the MCP server when the script is executed directly. | Launches the MCP server when the script is executed directly. | ||
== Dependencies == | |||
- | - <code>base64</code>: For encoding images | ||
- | - <code>cv2</code> (OpenCV): For video metadata extraction | ||
- | - <code>torch</code>: For tensor operations | ||
- | - <code>mcp.server.fastmcp</code>: FastMCP framework | ||
- | - <code>mcp.types</code>: MCP content types (ImageContent) | ||
- | - <code>qwen_vl_utils</code>: Video processing utilities (fetch_video) | ||
- | - <code>torchvision.transforms.functional</code>: Image conversion (to_pil_image) | ||
== Error Handling == | |||
=== Validation Errors === | |||
- Missing parameters: ValueError with specific parameter name | - Missing parameters: ValueError with specific parameter name | ||
- Invalid parameter values: ValueError with constraint details | - Invalid parameter values: ValueError with constraint details | ||
| Line 115: | Line 115: | ||
- Invalid video file: RuntimeError if file cannot be opened | - Invalid video file: RuntimeError if file cannot be opened | ||
=== Processing Errors === | |||
- Time range exceeds duration: ValueError with duration details | - Time range exceeds duration: ValueError with duration details | ||
- Video processing failure: RuntimeError with original exception context | - Video processing failure: RuntimeError with original exception context | ||
== Configuration == | |||
=== Video Processing Settings === | |||
- | - '''FPS''': 1 frame per second extraction rate | ||
- | - '''Min Frames''': 1 (minimum frames to extract) | ||
- | - '''Max Frames''': 128 (maximum frames to extract) | ||
- | - '''Max Pixels''': 224 × 224 (resolution constraint) | ||
=== Logging Settings === | |||
- | - '''Log Level''': INFO | ||
- | - '''Log File''': /tmp/mcp_server_debug.log | ||
- | - '''Console Output''': Enabled | ||
== Usage Example == | |||
=== Starting the Server === | |||
<syntaxhighlight lang="bash"> | |||
python examples/mcp_server/crop_video_mcp_server.py | python examples/mcp_server/crop_video_mcp_server.py | ||
</syntaxhighlight> | |||
=== Invoking from Client === | |||
<syntaxhighlight lang="python"> | |||
from lmms_eval.mcp.client import MCPClient | from lmms_eval.mcp.client import MCPClient | ||
| Line 157: | Line 157: | ||
# Convert to OpenAI format | # Convert to OpenAI format | ||
openai_content = client.convert_result_to_openai_format(result.content) | openai_content = client.convert_result_to_openai_format(result.content) | ||
</syntaxhighlight> | |||
== Design Decisions == | |||
1. | 1. '''1 FPS Extraction''': Balances detail with computational efficiency for video analysis | ||
2. | 2. '''Base64 Encoding''': Enables transmission of image data through text-based protocol | ||
3. | 3. '''PNG Format''': Lossless compression suitable for analysis tasks | ||
4. | 4. '''Comprehensive Validation''': Prevents errors early with clear messages | ||
5. | 5. '''Frame Limit''': 128 max frames prevents memory issues with long video segments | ||
== Related Components == | |||
- [[Sample_MCP_Server]]: Example of simpler MCP tools | - [[Sample_MCP_Server]]: Example of simpler MCP tools | ||
- [[MCP_Client]]: Client for invoking this server | - [[MCP_Client]]: Client for invoking this server | ||
- [[Media_Handling]]: Related video processing in main framework | - [[Media_Handling]]: Related video processing in main framework | ||
== Best Practices == | |||
1. Always validate time ranges before processing | 1. Always validate time ranges before processing | ||
2. Check video file accessibility and format | 2. Check video file accessibility and format | ||
Latest revision as of 10:38, 27 September 2026
File: examples/mcp_server/crop_video_mcp_server.py
Principle: MCP_Tool_Integration
Overview
The Crop Video MCP Server is a FastMCP-based server that provides video cropping functionality through the Model Context Protocol. It exposes a crop_video tool that extracts frames from a specified time range in a video file and returns them as a sequence of base64-encoded images.
Key Components
1. Server Initialization
app = FastMCP("Video Tools MCP Server", "0.1.0")
Creates a FastMCP application instance with name and version identifier.
2. Logging Configuration
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s [MCP_SERVER] %(levelname)s: %(message)s",
handlers=[logging.FileHandler("/tmp/mcp_server_debug.log"), logging.StreamHandler()],
)
Sets up dual logging to both file and console for debugging server operations.
3. Tool Implementation
crop_video Function
@app.tool(name="crop_video", description="Crop a video to a specified duration.")
def crop_video(
video_path: Annotated[str, Field(description="Path to the video file")] = None,
start_time: Annotated[float, Field(description="Start time in seconds")] = None,
end_time: Annotated[float, Field(description="End time in seconds, must be > start_time")] = None,
) -> list[ImageContent]:
Parameters:
- video_path (str): Path to the video file to process
- start_time (float): Start time in seconds for cropping
- end_time (float): End time in seconds (must be greater than start_time)
Returns: List of ImageContent objects containing extracted frames as base64 PNG images
Key Operations:
1. Parameter Validation:
- Checks all required parameters are provided - Validates parameter values (non-negative times, end > start) - Verifies video file existence
2. Video Duration Verification:
```python cap = cv2.VideoCapture(video_path) fps = cap.get(cv2.CAP_PROP_FPS) frame_count = cap.get(cv2.CAP_PROP_FRAME_COUNT) duration = frame_count / fps if fps > 0 else 0 ``` Uses OpenCV to check video duration and validate time range
3. Frame Extraction:
```python
video_ele = {
"type": "video",
"video": f"file://{video_path}",
"fps": 1, # 1fps
"min_frames": 1,
"max_frames": 128,
"max_pixels": 224 * 224,
"video_start": start_time,
"video_end": end_time,
}
video_frames = fetch_video(video_ele)
```
Uses qwen_vl_utils.fetch_video to extract frames at 1 FPS
4. Image Encoding:
```python video_frames = video_frames.to(torch.uint8) images = [to_pil_image(frame) for frame in video_frames]
image_contents = []
for img in images:
output_buffer = BytesIO()
img.save(output_buffer, format="PNG")
byte_data = output_buffer.getvalue()
base64_str = base64.b64encode(byte_data).decode("utf-8")
image_contents.append(ImageContent(type="image", data=base64_str, mimeType="image/png"))
```
Converts frames to PIL images, encodes as PNG, and wraps in ImageContent
4. Server Entry Point
if __name__ == "__main__":
app.run()
Launches the MCP server when the script is executed directly.
Dependencies
- base64: For encoding images
- cv2 (OpenCV): For video metadata extraction
- torch: For tensor operations
- mcp.server.fastmcp: FastMCP framework
- mcp.types: MCP content types (ImageContent)
- qwen_vl_utils: Video processing utilities (fetch_video)
- torchvision.transforms.functional: Image conversion (to_pil_image)
Error Handling
Validation Errors
- Missing parameters: ValueError with specific parameter name - Invalid parameter values: ValueError with constraint details - File not found: FileNotFoundError with file path - Invalid video file: RuntimeError if file cannot be opened
Processing Errors
- Time range exceeds duration: ValueError with duration details - Video processing failure: RuntimeError with original exception context
Configuration
Video Processing Settings
- FPS: 1 frame per second extraction rate - Min Frames: 1 (minimum frames to extract) - Max Frames: 128 (maximum frames to extract) - Max Pixels: 224 × 224 (resolution constraint)
Logging Settings
- Log Level: INFO - Log File: /tmp/mcp_server_debug.log - Console Output: Enabled
Usage Example
Starting the Server
python examples/mcp_server/crop_video_mcp_server.py
Invoking from Client
from lmms_eval.mcp.client import MCPClient
client = MCPClient("examples/mcp_server/crop_video_mcp_server.py")
# Get tool schema
functions = client.get_function_list_sync()
# Crop video from 5s to 10s
result = client.run_tool_sync("crop_video", {
"video_path": "/path/to/video.mp4",
"start_time": 5.0,
"end_time": 10.0
})
# Convert to OpenAI format
openai_content = client.convert_result_to_openai_format(result.content)
Design Decisions
1. 1 FPS Extraction: Balances detail with computational efficiency for video analysis 2. Base64 Encoding: Enables transmission of image data through text-based protocol 3. PNG Format: Lossless compression suitable for analysis tasks 4. Comprehensive Validation: Prevents errors early with clear messages 5. Frame Limit: 128 max frames prevents memory issues with long video segments
Related Components
- Sample_MCP_Server: Example of simpler MCP tools - MCP_Client: Client for invoking this server - Media_Handling: Related video processing in main framework
Best Practices
1. Always validate time ranges before processing 2. Check video file accessibility and format 3. Handle video processing exceptions with context 4. Log parameter values for debugging 5. Use appropriate frame rate for use case 6. Consider memory constraints with max frames 7. Provide clear error messages for validation failures