Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:EvolvingLMMs Lab Lmms eval Crop Video MCP Server

From Leeroopedia

File: examples/mcp_server/crop_video_mcp_server.py

Principle: MCP_Tool_Integration

Overview

The Crop Video MCP Server is a FastMCP-based server that provides video cropping functionality through the Model Context Protocol. It exposes a crop_video tool that extracts frames from a specified time range in a video file and returns them as a sequence of base64-encoded images.

Key Components

1. Server Initialization

app = FastMCP("Video Tools MCP Server", "0.1.0")

Creates a FastMCP application instance with name and version identifier.

2. Logging Configuration

logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s [MCP_SERVER] %(levelname)s: %(message)s",
    handlers=[logging.FileHandler("/tmp/mcp_server_debug.log"), logging.StreamHandler()],
)

Sets up dual logging to both file and console for debugging server operations.

3. Tool Implementation

crop_video Function

@app.tool(name="crop_video", description="Crop a video to a specified duration.")
def crop_video(
    video_path: Annotated[str, Field(description="Path to the video file")] = None,
    start_time: Annotated[float, Field(description="Start time in seconds")] = None,
    end_time: Annotated[float, Field(description="End time in seconds, must be > start_time")] = None,
) -> list[ImageContent]:

Parameters: - video_path (str): Path to the video file to process - start_time (float): Start time in seconds for cropping - end_time (float): End time in seconds (must be greater than start_time)

Returns: List of ImageContent objects containing extracted frames as base64 PNG images

Key Operations:

1. Parameter Validation:

  - Checks all required parameters are provided
  - Validates parameter values (non-negative times, end > start)
  - Verifies video file existence

2. Video Duration Verification:

  ```python
  cap = cv2.VideoCapture(video_path)
  fps = cap.get(cv2.CAP_PROP_FPS)
  frame_count = cap.get(cv2.CAP_PROP_FRAME_COUNT)
  duration = frame_count / fps if fps > 0 else 0
  ```
  Uses OpenCV to check video duration and validate time range

3. Frame Extraction:

  ```python
  video_ele = {
      "type": "video",
      "video": f"file://{video_path}",
      "fps": 1,  # 1fps
      "min_frames": 1,
      "max_frames": 128,
      "max_pixels": 224 * 224,
      "video_start": start_time,
      "video_end": end_time,
  }
  video_frames = fetch_video(video_ele)
  ```
  Uses qwen_vl_utils.fetch_video to extract frames at 1 FPS

4. Image Encoding:

  ```python
  video_frames = video_frames.to(torch.uint8)
  images = [to_pil_image(frame) for frame in video_frames]
  image_contents = []
  for img in images:
      output_buffer = BytesIO()
      img.save(output_buffer, format="PNG")
      byte_data = output_buffer.getvalue()
      base64_str = base64.b64encode(byte_data).decode("utf-8")
      image_contents.append(ImageContent(type="image", data=base64_str, mimeType="image/png"))
  ```
  Converts frames to PIL images, encodes as PNG, and wraps in ImageContent

4. Server Entry Point

if __name__ == "__main__":
    app.run()

Launches the MCP server when the script is executed directly.

Dependencies

- base64: For encoding images - cv2 (OpenCV): For video metadata extraction - torch: For tensor operations - mcp.server.fastmcp: FastMCP framework - mcp.types: MCP content types (ImageContent) - qwen_vl_utils: Video processing utilities (fetch_video) - torchvision.transforms.functional: Image conversion (to_pil_image)

Error Handling

Validation Errors

- Missing parameters: ValueError with specific parameter name - Invalid parameter values: ValueError with constraint details - File not found: FileNotFoundError with file path - Invalid video file: RuntimeError if file cannot be opened

Processing Errors

- Time range exceeds duration: ValueError with duration details - Video processing failure: RuntimeError with original exception context

Configuration

Video Processing Settings

- FPS: 1 frame per second extraction rate - Min Frames: 1 (minimum frames to extract) - Max Frames: 128 (maximum frames to extract) - Max Pixels: 224 × 224 (resolution constraint)

Logging Settings

- Log Level: INFO - Log File: /tmp/mcp_server_debug.log - Console Output: Enabled

Usage Example

Starting the Server

python examples/mcp_server/crop_video_mcp_server.py

Invoking from Client

from lmms_eval.mcp.client import MCPClient

client = MCPClient("examples/mcp_server/crop_video_mcp_server.py")

# Get tool schema
functions = client.get_function_list_sync()

# Crop video from 5s to 10s
result = client.run_tool_sync("crop_video", {
    "video_path": "/path/to/video.mp4",
    "start_time": 5.0,
    "end_time": 10.0
})

# Convert to OpenAI format
openai_content = client.convert_result_to_openai_format(result.content)

Design Decisions

1. 1 FPS Extraction: Balances detail with computational efficiency for video analysis 2. Base64 Encoding: Enables transmission of image data through text-based protocol 3. PNG Format: Lossless compression suitable for analysis tasks 4. Comprehensive Validation: Prevents errors early with clear messages 5. Frame Limit: 128 max frames prevents memory issues with long video segments

Related Components

- Sample_MCP_Server: Example of simpler MCP tools - MCP_Client: Client for invoking this server - Media_Handling: Related video processing in main framework

Best Practices

1. Always validate time ranges before processing 2. Check video file accessibility and format 3. Handle video processing exceptions with context 4. Log parameter values for debugging 5. Use appropriate frame rate for use case 6. Consider memory constraints with max frames 7. Provide clear error messages for validation failures

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment