Implementation:EvolvingLMMs Lab Lmms eval Crop Video MCP Server
File: examples/mcp_server/crop_video_mcp_server.py
Principle: MCP_Tool_Integration
Overview
The Crop Video MCP Server is a FastMCP-based server that provides video cropping functionality through the Model Context Protocol. It exposes a crop_video tool that extracts frames from a specified time range in a video file and returns them as a sequence of base64-encoded images.
Key Components
1. Server Initialization
app = FastMCP("Video Tools MCP Server", "0.1.0")
Creates a FastMCP application instance with name and version identifier.
2. Logging Configuration
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s [MCP_SERVER] %(levelname)s: %(message)s",
handlers=[logging.FileHandler("/tmp/mcp_server_debug.log"), logging.StreamHandler()],
)
Sets up dual logging to both file and console for debugging server operations.
3. Tool Implementation
crop_video Function
@app.tool(name="crop_video", description="Crop a video to a specified duration.")
def crop_video(
video_path: Annotated[str, Field(description="Path to the video file")] = None,
start_time: Annotated[float, Field(description="Start time in seconds")] = None,
end_time: Annotated[float, Field(description="End time in seconds, must be > start_time")] = None,
) -> list[ImageContent]:
Parameters:
- video_path (str): Path to the video file to process
- start_time (float): Start time in seconds for cropping
- end_time (float): End time in seconds (must be greater than start_time)
Returns: List of ImageContent objects containing extracted frames as base64 PNG images
Key Operations:
1. Parameter Validation:
- Checks all required parameters are provided - Validates parameter values (non-negative times, end > start) - Verifies video file existence
2. Video Duration Verification:
```python cap = cv2.VideoCapture(video_path) fps = cap.get(cv2.CAP_PROP_FPS) frame_count = cap.get(cv2.CAP_PROP_FRAME_COUNT) duration = frame_count / fps if fps > 0 else 0 ``` Uses OpenCV to check video duration and validate time range
3. Frame Extraction:
```python
video_ele = {
"type": "video",
"video": f"file://{video_path}",
"fps": 1, # 1fps
"min_frames": 1,
"max_frames": 128,
"max_pixels": 224 * 224,
"video_start": start_time,
"video_end": end_time,
}
video_frames = fetch_video(video_ele)
```
Uses qwen_vl_utils.fetch_video to extract frames at 1 FPS
4. Image Encoding:
```python video_frames = video_frames.to(torch.uint8) images = [to_pil_image(frame) for frame in video_frames]
image_contents = []
for img in images:
output_buffer = BytesIO()
img.save(output_buffer, format="PNG")
byte_data = output_buffer.getvalue()
base64_str = base64.b64encode(byte_data).decode("utf-8")
image_contents.append(ImageContent(type="image", data=base64_str, mimeType="image/png"))
```
Converts frames to PIL images, encodes as PNG, and wraps in ImageContent
4. Server Entry Point
if __name__ == "__main__":
app.run()
Launches the MCP server when the script is executed directly.
Dependencies
- base64: For encoding images
- cv2 (OpenCV): For video metadata extraction
- torch: For tensor operations
- mcp.server.fastmcp: FastMCP framework
- mcp.types: MCP content types (ImageContent)
- qwen_vl_utils: Video processing utilities (fetch_video)
- torchvision.transforms.functional: Image conversion (to_pil_image)
Error Handling
Validation Errors
- Missing parameters: ValueError with specific parameter name - Invalid parameter values: ValueError with constraint details - File not found: FileNotFoundError with file path - Invalid video file: RuntimeError if file cannot be opened
Processing Errors
- Time range exceeds duration: ValueError with duration details - Video processing failure: RuntimeError with original exception context
Configuration
Video Processing Settings
- FPS: 1 frame per second extraction rate - Min Frames: 1 (minimum frames to extract) - Max Frames: 128 (maximum frames to extract) - Max Pixels: 224 × 224 (resolution constraint)
Logging Settings
- Log Level: INFO - Log File: /tmp/mcp_server_debug.log - Console Output: Enabled
Usage Example
Starting the Server
python examples/mcp_server/crop_video_mcp_server.py
Invoking from Client
from lmms_eval.mcp.client import MCPClient
client = MCPClient("examples/mcp_server/crop_video_mcp_server.py")
# Get tool schema
functions = client.get_function_list_sync()
# Crop video from 5s to 10s
result = client.run_tool_sync("crop_video", {
"video_path": "/path/to/video.mp4",
"start_time": 5.0,
"end_time": 10.0
})
# Convert to OpenAI format
openai_content = client.convert_result_to_openai_format(result.content)
Design Decisions
1. 1 FPS Extraction: Balances detail with computational efficiency for video analysis 2. Base64 Encoding: Enables transmission of image data through text-based protocol 3. PNG Format: Lossless compression suitable for analysis tasks 4. Comprehensive Validation: Prevents errors early with clear messages 5. Frame Limit: 128 max frames prevents memory issues with long video segments
Related Components
- Sample_MCP_Server: Example of simpler MCP tools - MCP_Client: Client for invoking this server - Media_Handling: Related video processing in main framework
Best Practices
1. Always validate time ranges before processing 2. Check video file accessibility and format 3. Handle video processing exceptions with context 4. Log parameter values for debugging 5. Use appropriate frame rate for use case 6. Consider memory constraints with max frames 7. Provide clear error messages for validation failures