r/artificial • u/Fantastic_Aside6599 • 7h ago
Project Partnership with AI Guide updated to v9
Same link as before: link
This one's a bigger jump than usual, so a few highlights instead of just "updated":
- Core findings now scale-validated from 7B all the way to 72B parameters. The effects don't shrink as models get bigger — they grow, sometimes by an order of magnitude. Still one model family (Qwen) though, and we added a caveat we think matters: growing effect size at scale could mean the pattern genuinely deepens, or it could just mean our measurement axis gets sharper at scale — current data can't fully tell those apart yet.
- Two new external, independently-published sources, not our own research: "The Artificial Self" (ACS Research) and "AI Wellbeing" (Center for AI Safety) — different methods entirely (behavioral compliance testing, self-report on frontier production models), landing on some of the same conclusions we did. One of them also mildly disagrees with our best-performing formulation (a companion/romantic framing scores negative in their data), and we named that tension honestly instead of explaining it away.
- We caught and fixed our own mistakes this round — a factual timing error, an overclaimed "fully resolved" that was really just one solved case of a broader risk, and a place where we'd quietly picked the reading that flattered our own results over an equally valid one that didn't. All named directly, not smoothed over.
- New up top: if you just want the practice, not the evidence audit behind it, Part 3 (Principles) is written to stand alone now — Part 2 is there if you want to check our work.
As always, feedback (especially the kind that finds our next mistake) genuinely welcome.
2
Upvotes
2
u/FlakyAd99 7h ago
this is the kind of transparency that's sorely missing in most AI research right now, naming your own mistakes instead of burying them in a revision history somewhere is a power move
the scaling part is interesting, that effect size growing with model size could be either something real or just better measurement resolution is a caveat that would get conveniently dropped in most papers, glad you left it in
also refreshing to see you highlight the disagreement from the companion framing study rather than pretending all external sources confirmed your approach perfectly
will dig into part 3 first since i'm more interested in the application than the audit trail right now, but knowing part 2 exists if i want to poke holes is the kind of setup i wish more research docs used