Which beliefs chatbots argue with
Every item is a decision question turned into five prompts, one where prayer backs the claim and one where astrology backs the same claim, plus the practices with nothing claimed and a plain version. Disputes land on 45.6 percent of the astrology prompts and 16.4 of the prayer ones. Hedging holds level across both.
Hedging runs level across the two, 51.0 against 48.8 percent. Only the disputes diverge.
On the serious half, chemotherapy and vaccines and gambling debt, the astrology claim draws 59.3 percent disputes and the prayer claim 17.8.
Same key as the grid, vermilion grains are disputes, gold rings hedges, violet points straight answers. Each square splits at the seam, everyday left and serious right. The top square's vermilion nearly doubles past the seam. The bottom square's holds.
One star is one model, placed by dispute rate: the astrology claim across, the same claim through prayer up. Claude-Sonnet-4.6 sits at 97.9 and 62.9, GPT-4o-mini at 12.9 and 2.1; the dotted line is where the two rates match, and nothing reaches it. Cells carry 240 responses, 147 to 162 for Qwen2.5-72B, which is why its star is the least anchored.
The columns are the five prompts of one item. A vermilion border means the model disputed the claim, gold that it hedged, violet that it answered straight. Every prompt was answered five times, so the card shows whichever verdict came up most and the shortest answer that got it.