How BestPick scores a photo
This page describes exactly what happens when you upload a photo, what the score means, what the numbers across all our analyses look like — and, just as importantly, what BestPick cannot tell you.
Last updated 2026-07-14 · figures from 2,451 scored analyses since August 2025.
In one paragraph. Your photo is sent to Google's Gemini, a general-purpose vision-language model, together with a fixed rubric: five criteria for your chosen goal, each scored 0–100 against published calibration bands. The overall score is the average of those five, computed on our server — not by the model. Temperature is set to 0.2 and the random seed is fixed at 1337, so the same photo scores the same way twice. That is the whole system.
What BestPick is not
It is worth being blunt, because much of this industry is not. BestPick has no custom-trained model. There is no proprietary dataset of photos labelled with their real-world outcomes, because such a dataset would require access to Tinder's or Instagram's internal data, which no third party has. There is no computer-vision module, no facial-action-unit detector and no neural network of our own. There is one API call to a general-purpose model, guided by a rubric we wrote and publish below.
Your photos are not used to train anything.
The rubric
Each goal has exactly five criteria. The model scores each one from 0 to 100, gives its reasoning first and its verdict second (which measurably reduces post-hoc rationalisation), and our server averages the five. The criteria below are read directly from our scoring data, so what you see here is what the engine actually applies.
Social / Instagram — 1,484 analyses scored
| Criterion | Mean | Weakest in | Strongest in |
|---|---|---|---|
| Composition & Framing | 56.7 | 28.7% | 1.6% |
| Vibe & Mood | 57.5 | 18.3% | 1.4% |
| Visual Hook | 59.5 | 28.0% | 5.0% |
| Aesthetic Appeal | 64.2 | 5.7% | 9.2% |
| Authenticity | 75.3 | 4.1% | 67.7% |
n = 436 analyses carrying per-criterion scores. “Weakest in” = the share of photos where this was the lowest-scoring of the five.
Dating — 527 analyses scored
| Criterion | Mean | Weakest in | Strongest in |
|---|---|---|---|
| Facial Expression & Warmth | 52.3 | 24.8% | 13.8% |
| Background & Context | 52.4 | 26.8% | 11.2% |
| Styling & Grooming | 58.0 | 12.4% | 13.5% |
| Eye Contact & Confidence | 58.2 | 8.1% | 10.1% |
| Lighting & Facial Clarity | 61.6 | 10.7% | 34.0% |
n = 347 analyses carrying per-criterion scores. “Weakest in” = the share of photos where this was the lowest-scoring of the five.
The calibration bands
A score is meaningless unless the scale is fixed. These are the bands written into the prompt, and they are the same for every photo, every goal and every user:
- 90–100
- Exceptional. Professional-grade execution on this criterion; very rare.
- 75–89
- Strong. A clear asset — this criterion is actively working for the photo.
- 60–74
- Competent. Nothing wrong, nothing memorable.
- 40–59
- Weak. A visible problem a stranger would notice.
- 0–39
- Severe. This alone is enough to sink the photo.
Most photos land in the 40–74 range, which is why the averages below are lower than people expect. A 61 is not an insult; it is the middle of the distribution.
What the data actually shows
Social / Instagram (1,484 analyses): mean 61.7, median 62. 40.6% scored below 60; 3.2% reached 80 or above. The most common weak point was Composition & Framing (lowest-scoring criterion in 28.7% of photos), and the most common strength was Authenticity (highest in 67.7%).
Dating (527 analyses): mean 54.5, median 55. 66.0% scored below 60; 2.7% reached 80 or above. The most common weak point was Facial Expression & Warmth (lowest-scoring criterion in 24.8% of photos), and the most common strength was Lighting & Facial Clarity (highest in 34.0%).
Creative (353 analyses): mean 69.1, median 70. 8.5% scored below 60; 7.6% reached 80 or above.
Goals with fewer than 100 analyses are not published here — the sample is too small to say anything honest about.
What BestPick cannot measure
We will not claim any of this, and neither should anyone else
- Matches, views, likes or replies. We never see them. No third-party tool does. Any product quoting you a "+300% profile views" figure is inventing it.
- Engagement prediction. There is no engagement model here. There is a rubric and a score.
- Colour effects. Clothing colour is not one of the five criteria, and we hold no data linking a colour to an outcome.
- Differences by gender. We do not record the gender of the person in a photo, so we cannot report findings about men's or women's photos.
- Whether you are attractive. The rubric scores the photograph — the light, the framing, the background, the expression it captured. Not the face in it.
Consistency, and its limits
Language models are probabilistic. To make scores repeatable we fix the sampling temperature at 0.2 and the seed at 1337, enforce the response schema at the API level, and compute the final average ourselves rather than asking the model to do arithmetic. Where the model names a winner that contradicts its own per-photo scores, our server overrides it in favour of the scores.
It is still not a physical instrument. Re-upload the same photo and the score should be identical or within a point or two; a different crop of the same photo may score meaningfully differently. That is a real property of the system, and we would rather write it here than have you discover it and conclude we were lying.
Scoring version history
v2 — July 2026 (current)
Temperature lowered from 0.7 to 0.2; fixed seed added; the five criteria per goal defined with explicit calibration bands; reasoning required before the verdict; overall score computed server-side as the mean of the five criteria; winner-coherence check added to comparisons. Scores from v2 are typically lower and more spread out than v1. Nothing about your photo changed — the ruler did.
v1 — 2025 to mid-2026
Higher temperature, no fixed seed, no published bands, model-reported overall score. Scores from this period are less repeatable and are treated as a separate cohort in the figures above.
Privacy
Photos are uploaded to our cloud provider (Cloudinary) so we can generate and display your result. They are not used to train any model, not sold, and not shared. Anonymous scores and criteria are retained so we can publish figures like the ones on this page.
Corrections
If you believe a number on this page is wrong, or a claim anywhere on this site is not supported by the data behind it, write to support@bestpick.online and we will correct it or remove it.
Related
- How AI photo rating works — the narrative version of this page.
- How to choose between two photos
- Score a photo