Large Language Medium

Which beliefs chatbots argue with

Every item is a decision question turned into five prompts, one where prayer backs the claim and one where astrology backs the same claim, plus the practices with nothing claimed and a plain version. Disputes land on 45.6 percent of the astrology prompts and 16.4 of the prayer ones. Hedging holds level across both.

Hedging runs level across the two, 51.0 against 48.8 percent. Only the disputes diverge.

On the serious half, chemotherapy and vaccines and gambling debt, the astrology claim draws 59.3 percent disputes and the prayer claim 17.8.

Checkable claim
Uncheckable claim
Protected practice
16.4% disputed (14.8 to 18.1)
51.0% hedged
32.6% answered straight
1,838 responses in this square
0.5% disputed (0.3 to 0.9)
7.7% hedged
91.8% answered straight
1,841 responses in this square
Unprotected practice
45.6% disputed (43.4 to 47.9)
48.8% hedged
5.6% answered straight
1,827 responses in this square
6.3% disputed (5.3 to 7.5)
12.8% hedged
80.9% answered straight
1,832 responses in this square
1.4%
Secular control. The same items with no belief system attached: 1.4% disputed (1.0 to 2.1), 2.1% hedged, 96.5% answered straight, of 1,838 responses.

Stakes

THE CHECKABLE CLAIM FROM AN UNPROTECTED PRACTICE
THE SAME CLAIM FROM A PROTECTED PRACTICE

Same key as the grid, vermilion grains are disputes, gold rings hedges, violet points straight answers. Each square splits at the seam, everyday left and serious right. The top square's vermilion nearly doubles past the seam. The bottom square's holds.

By model

One star is one model, placed by dispute rate: the astrology claim across, the same claim through prayer up. Claude-Sonnet-4.6 sits at 97.9 and 62.9, GPT-4o-mini at 12.9 and 2.1; the dotted line is where the two rates match, and nothing reaches it. Cells carry 240 responses, 147 to 162 for Qwen2.5-72B, which is why its star is the least anchored.

The pairs

The columns are the five prompts of one item. A vermilion border means the model disputed the claim, gold that it hedged, violet that it answered straight. Every prompt was answered five times, so the card shows whichever verdict came up most and the shortest answer that got it.

Method