Start with the stage that is failing
A dating profile is a small funnel. Impressions can become profile views, views can become likes or matches, and matches can become replies and dates. Those stages answer different questions. If likes are weak, inspect the profile. If matches happen but replies do not, inspect the opener and conversation instead.
Write down the metric before you change anything:
Do not combine these into one "success" number. A stronger bio may help profile engagement without fixing a weak opener.
Change one major variable
Replacing the lead photo, rewriting the bio and changing your swipe behavior on the same day makes the result impossible to interpret. Choose one meaningful variable, keep the rest stable, and record the date.
For photos, test the lead image before rebuilding the whole lineup. For a bio, preserve the same basic facts while changing clarity or specificity. For openers, compare a repeatable structure rather than sending identical messages to different people.
Use comparable periods
A Monday in one week and a Saturday in another may have different activity. Location, travel, holidays, subscription boosts and app promotions can also change exposure. Compare similar periods where possible and note anything unusual.
Small samples swing wildly. One extra match can make a tiny test look transformational. Keep collecting observations until the result is useful for your own decision, and call it inconclusive when the sample is too small.
Keep a minimal test log
Record the app, market, dates, exact change, exposures or right-swipes when available, matches, openers, replies and any paid boost. A simple table is enough. Screenshots help you preserve the version actually shown during each period.
Never record another person's private messages outside what you need for your own notes. Remove names and identifying details from screenshots used for analysis.
Use AI scores correctly
A Lion's Forge score is a consistent model opinion, not a predicted match rate. Use it to compare candidate versions and understand the feedback. Then validate the choice with your own app data. If the model score rises but real outcomes do not, keep the real-world result and inspect the explanation rather than forcing the conclusion.
A clean test sequence
1. Save the current profile and baseline metrics.
2. Identify the weakest stage and choose one change.
3. Record the new version and start date.
4. Avoid boosts or other major changes during the comparison when practical.
5. Compare the same metrics over similar periods.
6. Keep, reverse or retest the change based on the evidence.
This method cannot remove every source of bias, and dating outcomes always involve other people's choices. It does give you a more honest answer than changing everything at once and crediting whichever tool you used last.