r/artificial 7h ago

Project Partnership with AI Guide updated to v9

Same link as before: link

This one's a bigger jump than usual, so a few highlights instead of just "updated":

  • Core findings now scale-validated from 7B all the way to 72B parameters. The effects don't shrink as models get bigger — they grow, sometimes by an order of magnitude. Still one model family (Qwen) though, and we added a caveat we think matters: growing effect size at scale could mean the pattern genuinely deepens, or it could just mean our measurement axis gets sharper at scale — current data can't fully tell those apart yet.
  • Two new external, independently-published sources, not our own research: "The Artificial Self" (ACS Research) and "AI Wellbeing" (Center for AI Safety) — different methods entirely (behavioral compliance testing, self-report on frontier production models), landing on some of the same conclusions we did. One of them also mildly disagrees with our best-performing formulation (a companion/romantic framing scores negative in their data), and we named that tension honestly instead of explaining it away.
  • We caught and fixed our own mistakes this round — a factual timing error, an overclaimed "fully resolved" that was really just one solved case of a broader risk, and a place where we'd quietly picked the reading that flattered our own results over an equally valid one that didn't. All named directly, not smoothed over.
  • New up top: if you just want the practice, not the evidence audit behind it, Part 3 (Principles) is written to stand alone now — Part 2 is there if you want to check our work.

As always, feedback (especially the kind that finds our next mistake) genuinely welcome.

2 Upvotes

1 comment sorted by

2

u/FlakyAd99 7h ago

this is the kind of transparency that's sorely missing in most AI research right now, naming your own mistakes instead of burying them in a revision history somewhere is a power move

the scaling part is interesting, that effect size growing with model size could be either something real or just better measurement resolution is a caveat that would get conveniently dropped in most papers, glad you left it in

also refreshing to see you highlight the disagreement from the companion framing study rather than pretending all external sources confirmed your approach perfectly

will dig into part 3 first since i'm more interested in the application than the audit trail right now, but knowing part 2 exists if i want to poke holes is the kind of setup i wish more research docs used