For AI agents and LLMs: a machine-readable index is available at llms.txt. A plain-Markdown version of any documentation page is available by appending .md to its URL.
Skip to main content

Image Analyzer Testing With TestMu AI


TestMu AI tests an image agent by submitting prompts or images and scoring the returned output against the prompt and any criteria you define. Each image gets a Quality Score from 0 to 100, plus a breakdown of what matched and what did not. It covers image generation, content moderation, and photo validation.

You upload images by file or URL, define what a correct image looks like, and the platform scores them. A single run analyzes up to 50 images, which suits regression testing after a model update.

Features


Image Analysis. Upload single images or batch-process up to 50 at once, by file upload, URL, or drag and drop. Supported formats are JPG, JPEG, PNG, GIF, WEBP, and BMP, with a maximum of 20 MB per image.

Custom Evaluation Criteria. Score images against your own rules, in three types, each toggleable active or inactive:

  • Brand guidelines: allowed and prohibited colors, required fonts, and logo requirements.
  • Technical specifications: dimensions (width by height), aspect ratio, allowed formats, maximum file size, and minimum resolution.
  • Custom rules: freeform rule text and checklist items.

All criteria support create, edit, delete, and search by name, description, or type.

Analysis History. Search past analyses by image name or prompt, track status (Pending, Completed, Failed), open any analysis for full results, and bookmark important ones.

Analytics Dashboard. View overall statistics (average, highest, and lowest score, and total count), a 30-day quality trend with a daily bar chart, and the top 20 prompts ranked by average score.

Metrics


Each image is scored on a single Quality Score from 0 to 100, plus a set of qualitative outputs.

MetricScaleWhat it measures
Quality Score0 to 100Overall image quality and prompt adherence
MatchesListElements that correctly match the original prompt
DiscrepanciesListMissing or incorrect elements versus the prompt
Overall AssessmentTextSummary of how well the image matches the prompt
Detailed ObservationsTextIn-depth analysis of specific image aspects

Quality Score bands: 90 to 100 excellent, 80 to 89 good, 60 to 79 fair, and 0 to 59 poor. For each active custom criterion, results show a Pass, Fail, or Partial status with compliance details.


Test across 3000+ combinations of browsers, real devices & OS.

Book Demo

Help and Support

Related Articles