01 — Research question
Can a multimodal model reproduce a UX first impression?
The public benchmark combines full-page homepage screenshots with human ratings. Its dimensions cover aesthetics, perceived usability, trustworthiness, typicality, family resemblance and exemplar goodness.
The experiment tests one measurable capability: recovering the relative ordering of human judgments on websites that were not designed for this benchmark. It does not measure actual use after interaction and therefore does not replace UX research or A/B testing.
02 — Experimental pipeline
Items, scale and screenshots remain aligned with the reference study.
Each screenshot is sent to the model with the six original items and anchors. The model returns one continuous score between −3 and +3 plus a structured qualitative rationale. An incremental cache avoids recomputing completed evaluations.
The reported experiment covers 100 fashion websites, producing 600 image-dimension pairs. Outputs are stored in Parquet and joined with human scores through each screenshot identifier, enabling analysis of relative ranking and absolute error.
03 — Results
The ranking contains signal; calibration remains insufficient.
Overall Pearson correlation reaches 0.416. It ranges from 0.343 for aesthetics to 0.535 for exemplar goodness, with 0.486 for usability and 0.513 for family resemblance.
The model nevertheless assigns overly positive values: dimension means range approximately from +1.60 to +2.15 while standardized human ratings remain centered near zero. Overall mean absolute error reaches 1.98 points on the six-point scale. The model ranks better than it calibrates.
04 — Decision
Using the model as a comparison instrument requires explicit guardrails.
The result supports exploration as a shortlisting and research tool: comparing variants, finding disagreements and generating hypotheses to validate with users. It does not support automatic judgments of design quality.
The next methodological steps include calibration on a held-out sample, replication across homeware, banking and higher-education websites, uncertainty intervals, and comparison with decisions and behavior observed in actual product experiments.