Elle Liemandt · 2025 · Randomized within-subjects repeated-measures experiment · n=404

An Experimental and Qualitative Comparison of the Quality of Dating Advice from Parents Versus LLMs

Cited 0 times in the scientific literature.

Level 2 - randomized trial

Randomized within-subjects experimental trial evaluated by design analogy to clinical CEBM standards.

OpenAlex W4417278746 · doi:10.31234/osf.io/r7wy4_v2 · record verified 2026-08-31

What was done

In a preregistered, within-subjects, repeated-measures randomized experiment, 404 adolescents aged 13 to 17 read three realistic dating conflict scenarios sourced from a relationship advice application. For each scenario, participants received advice counterbalanced and randomly attributed to three sources: parents, an off-the-shelf LLM (ChatGPT or Gemini), or a fine-tuned relationship advice LLM (AskElle). Participants rated each response for helpfulness with coping, likelihood of implementation, and trustworthiness. Data were analyzed using a random-intercept Bayesian machine-learning model alongside qualitative natural-language analysis.

What was found

LLM-generated advice outperformed parental advice by approximately 0.3 standard deviations across all primary outcomes (all pr(ATE > 0) > .99). Overall, the fine-tuned LLM performed comparably to off-the-shelf LLMs across trials. Off-the-shelf LLMs were initially preferred on the first presentation (pr(Diff) > .99), but participants strongly preferred the fine-tuned LLM on the third exposure and after direct comparison to other sources (pr(Diff) > .99 over parents and generic LLMs). Qualitative analysis revealed parents often minimized problems as temporary, generic LLMs were overly accommodating and potentially reinforced rumination, and the highest-rated answers validated emotions while challenging unhelpful appraisals and providing concrete scripts.

Why it matters

This study shows that adolescents find AI-generated relationship guidance more trustworthy and actionable than advice from parents. It also demonstrates that specialized, fine-tuned models outperform generic LLMs once users have repeated exposure and can compare outputs.

Limits

The study evaluated subjective ratings of hypothetical scenarios rather than real-time guidance during live interpersonal conflicts. Self-reported ratings of trustworthiness and intent do not measure actual behavioral execution, safety, or long-term relationship outcomes. Details regarding how parental advice was collected or standardized are not detailed in the abstract.

Cited by