Analyze, describe, and extract information from images using the MiniMax vision MCP tool. Use when: user shares an image file path or URL (any message containing .jpg, .jpeg, .png, .gif, .webp, .bmp, or .svg file extension) or uses any of these words/phrases near an image: "analyze", "analyse", "describe", "explain", "understand", "look at", "review", "extract text", "OCR", "what is in", "what's in", "read this image", "see this image", "tell me about", "explain this", "interpret this", in conne
Detected risks:
The vision-analysis skill enables AI agents to analyze, describe, and extract information from images using the MiniMax vision capabilities through the MCP (Model Context Protocol) tool integration. It automatically triggers when users share image files or request image analysis, providing intelligent interpretation of visual content including screenshots, diagrams, charts, UI mockups, and photographs. The skill leverages the MiniMax Token Plan's understand_image tool to deliver comprehensive visual understanding.
This skill supports multiple analysis modes tailored to different use cases: general description for overall image understanding, OCR for text extraction from documents and screenshots, UI review for design critique and feedback on mockups and wireframes, chart data extraction for analyzing graphs and visualizations, and object detection for identifying elements within images. Each mode uses optimized prompting strategies to deliver relevant results based on the specific analysis requirement.
Developers, designers, data analysts, and content creators can use this skill to automate image analysis tasks within their AI-assisted workflows. Common applications include extracting text from screenshots for documentation, reviewing UI designs for accessibility and usability issues, analyzing charts to extract numerical data, identifying objects or activities in photos, and generating detailed descriptions of visual content for reports or accessibility purposes. The skill requires a MiniMax Token Plan subscription and proper MCP configuration to function.
The skill supports common image formats including JPG, JPEG, PNG, GIF, WebP, BMP, and SVG files. It can analyze images provided as local file paths or URLs.
Yes, this skill requires an active MiniMax Token Plan subscription with a valid MINIMAX_API_KEY. The understand_image tool cannot be used with free or other tier API keys.
Configuration depends on your environment. For OpenCode, add the MCP configuration to opencode.json. For Claude Code, use the 'claude mcp add' command. For Cursor, add to MCP settings. After configuration, restart your application and verify with the /mcp command. Detailed setup instructions are available at https://platform.minimaxi.com/docs/token-plan/mcp-guide.
Five modes are available: 'describe' for general image understanding, 'ocr' for text extraction, 'ui-review' for design critique of mockups and wireframes, 'chart-data' for extracting information from graphs and visualizations, and 'object-detect' for identifying and locating elements within images.
The skill activates automatically when a message contains an image file extension or when users use trigger phrases like 'analyze', 'describe', 'explain', 'extract text', 'OCR', 'what is in', or 'review' in connection with an image, screenshot, diagram, chart, mockup, or photo.
Quick Setup:
.claude/skills/Repository
minimax-ai/skills