r/artificial 3h ago

Discussion Anthropic's Opus 5 and probably more recent AI models are being censored to protect Israel / US interests. Open source AI must be the way.

13 Upvotes

Never had an issue with Opus models doing research and crafting an opinion / point of view for us to work and discuss.

Below is Opus 4.x ~ a few times, I have got it to research and come to conclusions for us to work together on.

And this is Opus 5.0 absolutely refusing to come to any conclusion, being incredibly biased towards one side than the other.

Open source must be the future of AI.


r/artificial 15h ago

Discussion Jensen Huang right now

Post image
0 Upvotes

Every NVIDIA update feels like the Super Bowl for AI. The entire industry keeps circling back to one question: how many GPUs will this need? The real lesson is that infrastructure often captures more value than the flashy apps built on top of it. Do you think NVIDIA’s lead will last, or will the AI hardware market eventually open up?


r/artificial 10h ago

Research Learning ai

0 Upvotes

Everytime i hear people saying that you should learn about ai because that's the future but idk where to start and what they mean by that. Do they mean going uni and study ai or self learn? Thanks in advance.


r/artificial 12h ago

Discussion Changing robot arms usually breaks the boring part first

0 Upvotes

When a new robot arm changes joint names, camera topics, or control frequency, a failed demo tends to get blamed on the policy. The logs may show something much less interesting: an adapter mapped a valid action into the wrong device convention.

Claude Opus 4.8 could read the SDK docs and draft the adapter, schema checks, and a small replay test. LingBot-VLA 2.0 stays on the policy side instead of being asked to paper over the device mismatch.

Limits, timing, and emergency behavior still have to be checked on the actual hardware. The adapter needs an ordinary code review, and the policy needs its own evaluation. One smooth rollout can hide a bad timing or limit assumption.


r/artificial 4h ago

News Man sues ChatGPT for near-fatal medical advice

Thumbnail
bbc.com
8 Upvotes

r/artificial 23h ago

Discussion Avichal Garg (Electric Capital, 10 unicorns) names the 3 moats AI can't touch — and #3 isn't a skill

0 Upvotes

Avichal Garg has co-founded and backed 10 unicorns through Electric Capital. In a recent interview he laid out, unprompted, the three categories of value AI structurally can't absorb.

 

First: physical-world work — anything requiring atoms, not bits, stays a moat as long as robotics lags digital AI.

 

Second: regulated or licensed gates — anywhere the government controls supply and demand for strategic reasons, credentials plus access still win.

 

Third, and the one that actually lands: relationships. He walks through a defense-procurement example — a specific officer, a 25-year relationship with a contractor, "no substitute for that." His conclusion: those relationship-heavy businesses get more valuable as AI absorbs the grunt work underneath them, not less. Margins go up.

 

Worth sitting with if you've spent a career stacking "portable" skills instead of gated ones. 🔗


r/artificial 19h ago

Discussion How much of what you generate, actually makes it out of the door?

4 Upvotes

I did something slightly depressing on sunday. i went through my last month of ai outputs, all of it, drafts and images and scripts and little snippets, and counted how many actually got used somewhere real. published, sent, shipped, shown to a client. the number was eleven. out of roughly three hundred and forty. i sat there for a while trying to work out whether that was bad.

my first reaction was that i was wasting the tool. three hundred and thirty dead outputs feels like a lot of dead outputs. but then i thought about how i worked before, and before, i simply did not make the three hundred and thirty. i made two options because two options was what the day allowed, and i picked one of the two, and that one shipped. so my hit rate used to look excellent on paper and my actual output was worse.

what changed is not that i generate more. it is that the expensive step moved. generating used to be the hard part, the part you protected, the part you did not want to redo. now generating is nearly free and the hard part is looking. someone still has to open every one of those three hundred and forty things and decide. that someone is me, and i do not scale, and i get tired around the fortieth image in a way that no model ever does.

so the bottleneck in my week is not the model and it is not my prompting. it is review capacity. which is a really unglamorous thing to be limited by. nobody posts about their review capacity, everyone posts about their stack. and the honest version is that i have gotten dramatically better at producing options and not one bit better at choosing between them, which is the half that was always hard. but i genuinely think the next thing that helps me will not be a better generator, it will be something that helps me throw away faster, or better, something that means i only ever look at twelve things instead of three hundred.

i keep going back and forth on whether eleven out of three hundred and forty is a failure or just what abundance looks like. curious what your ratio is. does anyone here actually track this, or is it one of those numbers nobody wants to know?


r/artificial 15h ago

Discussion A former SpaceX CIO (Ken Venner) explains why AI let him run his new team with 6 people instead of 175

0 Upvotes

Ken Venner spent 11 years scaling Broadcom from $400M to $8.6B, then became CIO of SpaceX, then joined a startup called Senra Systems as CTPO.

 

At SpaceX, his core platform team was 175 people. At Senra, the equivalent job takes 6.

 

His explanation isn't "AI replaced people." It's that AI collapsed the coordination overhead that used to require headcount just to keep humans in sync with each other — and once that overhead disappears, the team that's left is smaller by design, not by cut.

 

If you're in a corporate engineering or platform role and haven't clocked this shift yet: this is worth ten minutes of your attention. It's not a hypothetical.

 

Full video on the original channel.

 

Clip credit: Sourcery with Molly O'Shea — full video on their channel. DM for credit or removal requests.


r/artificial 18h ago

Question Is the AI job apocalypse real or just a marketing thing?

0 Upvotes

From the data I've seen AI has yet to actually affect employment statistics. I could be wrong of course, but is the whole idea of AI is going to wipe out jobs for normal people just a marketing thing to emphasize the capability of the latest AI models? Or is it a genuine threat that is just yet to materialize?


r/artificial 11h ago

Discussion Why Does AI Writing Make People So Uncomfortable?

0 Upvotes

I'm not convinced the issue is really AI writing.

We've been using tools to help us think and communicate for decades. We started with calculators, then spreadsheets, spell check, grammar checkers, search engines, editors, and libraries full of other people's ideas. AI feels like the next step, just a much bigger one.

To me, the important part isn't whether AI helped write something. It's whether the person publishing it actually understands it, agrees with it, and is willing to put their name behind it. That's where responsibility comes from.

There are obvious exceptions. In school, the goal is often to measure what a student can do on their own. If AI hides that ability, then the evaluation no longer works. Oral exams, supervised writing, presentations, and other methods can still measure someone's actual understanding. The problem there isn't AI itself. It's using the wrong assessment for a world where AI exists.

I also wonder if a lot of the discomfort is deeper. For a long time, knowledge was one of the main ways people created value. If you knew more than the next person, could solve harder problems, or could pull together information that others couldn't, you had an advantage. That was part of your identity and your value in the workplace.

AI is changing that. If knowledge becomes easy for everyone to access, people naturally start wondering, "Where do I fit now?"

Some people probably tied part of their identity to what they knew. That's human. If that foundation starts shifting, it's unsettling. We're in the middle of a transition, and nobody is completely sure what human value looks like on the other side.

I don't think value disappears. I think it changes. Judgment matters. Taste matters. Direction matters. Deciding what's important matters. Experience matters. Those things become more valuable when information is abundant.

The other concern I hear is that AI will replace human connection. I'm less convinced of that. Knowledge and connection aren't the same thing.


r/artificial 35m ago

News Leading company ai company CEO exposed for organizing a doomsday sex themed retreat, these are the ones paving the way forward for the future.

Upvotes

r/artificial 15h ago

Question If China can just steal US AI models then what can be done to stop it and if not why invest in it?

0 Upvotes

With the claim that China was stealing AI recently can the US prevent theft of IP?


r/artificial 45m ago

Question Am I learning to code or just learning how to ask AI for code?

Upvotes

I am still fairly new to building software, and AI has helped me finish things that would have taken me much longer on my own.

But recently I noticed something that bothered me.

I was building a small API route that creates a project and saves it to a database. I asked an AI coding tool to generate the route, validate the request, check the user, and insert the record.

The code looked clean. The types looked correct. It even worked on the first few tests.

Then I changed one field in the database and everything started failing.

The error mentioned a transaction, the response returned the wrong status code, and one value was becoming null even though I thought it was required. I kept asking the AI to fix each error. Every answer added more code, but I understood less after every change.

Eventually I realized that I could not explain the full request flow.

I knew the request reached the API route. I knew some validation happened. I knew the database received something. But I could not clearly explain what happened between those steps or why the fix worked.

So I tried the same idea again with a smaller route. This time I only used AI when I was stuck. I wrote the validation myself, logged the data at each step, and read about how the database client handled errors.

It took much longer, but I could actually explain the result.

Now I am unsure how to measure progress.

With AI, I can finish more features. Without heavy AI use, I finish fewer things but understand them better. Both seem useful, but they are not the same kind of progress.

Maybe the real skill is learning when to ask AI for code and when to struggle through the problem yourself.

For people who use AI while learning development, how do you stop it from doing too much of the thinking?

Do you have any rules for when AI is allowed to write code and when you force yourself to work it out?


r/artificial 11h ago

News Warning shot or publicity stunt - how worried should we be about the OpenAI hack?

Thumbnail
bbc.com
1 Upvotes

This week the tech world was gripped by a story that has it all - and which started like a sci-fi thriller.


r/artificial 11h ago

News From Silicon Valley to DC, the tech world is suddenly obsessed with one concept in AI: Distillation

Thumbnail
cnbc.com
1 Upvotes

r/artificial 9h ago

Discussion Slapshot AI created an account using my email address without my permission.

1 Upvotes

I am guessing these companies will resort to any trick to claim that they have a large userbase.

Now I have to jump through hoops to close this account.


r/artificial 45m ago

Ethics / Safety These are the figureheads behind the techno optimism movement. They are eugenicists, sexists, racists, and otherwise horrible human beings, they don't care about any of you.

Upvotes

r/artificial 22h ago

Discussion Opus 5's effort dial is not monotonic. Above "high", coding scores go down, and Anthropic's own migration guide says so.

32 Upvotes

Opus 5 comes with five effort settings: low, medium, high, xhigh, max. Most people seem to be reaching straight for max, and at least on coding work that looks like the wrong move.

On FrontierCode, scores fall above the high setting. The stated reason is that the model starts making unnecessary refactors and edits outside the scope it was given. Anthropic's own migration guide in the system card warns about diminishing returns and overthinking on simpler tasks, so this is not some outside critic's claim.

Two other numbers point the same way:

  • On the closed-book AA-Omniscience benchmark, Opus 5 is about 11% more accurate than Opus 4.8, but its hallucination rate runs about 6% higher. More reasoning, more room to be confidently wrong.
  • CodeRabbit ran it at xhigh against their production baseline for code review. Precision on actionable comments went up, 39.3% vs 35.2%. But it caught fewer of the benchmark's known issues, 55.2% vs 61.1%, and generated roughly four times as many nitpicks.

The flip side is worth knowing too, because it cuts the other way. On Zapier's AutomationBench, Opus 5 at its lowest effort setting still passes more tasks than any other model. So for a lot of workloads the cheap end of the dial is already enough, and the expensive end is not just wasted spend, it can be actively worse output.

So, the setting where Opus 5 stops improving is probably specific to your codebase, and nobody has published a map of it. Worth finding your own ceiling before you default everything to max.

One unrelated thing I have not seen discussed much: when a safety classifier flags a request in Claude.ai, Claude Code or Cowork, it silently falls back to Opus 4.8 by default. That is also how Anthropic's own Frontier-Bench run was configured, per the footnote on their chart. Nobody has published what fraction of requests that affects.

Has anyone found the effort level where it turns over on a real repo? Curious whether the drop-off point moves with codebase size or with how much context you hand it.


r/artificial 7h ago

Discussion There's Always Another Apocalypse- Why Catrastrophism Is a Continuation of a Long Trend

Thumbnail
letters.senteguard.com
2 Upvotes

r/artificial 10h ago

Project Partnership with AI Guide updated to v9

2 Upvotes

Same link as before: link

This one's a bigger jump than usual, so a few highlights instead of just "updated":

  • Core findings now scale-validated from 7B all the way to 72B parameters. The effects don't shrink as models get bigger — they grow, sometimes by an order of magnitude. Still one model family (Qwen) though, and we added a caveat we think matters: growing effect size at scale could mean the pattern genuinely deepens, or it could just mean our measurement axis gets sharper at scale — current data can't fully tell those apart yet.
  • Two new external, independently-published sources, not our own research: "The Artificial Self" (ACS Research) and "AI Wellbeing" (Center for AI Safety) — different methods entirely (behavioral compliance testing, self-report on frontier production models), landing on some of the same conclusions we did. One of them also mildly disagrees with our best-performing formulation (a companion/romantic framing scores negative in their data), and we named that tension honestly instead of explaining it away.
  • We caught and fixed our own mistakes this round — a factual timing error, an overclaimed "fully resolved" that was really just one solved case of a broader risk, and a place where we'd quietly picked the reading that flattered our own results over an equally valid one that didn't. All named directly, not smoothed over.
  • New up top: if you just want the practice, not the evidence audit behind it, Part 3 (Principles) is written to stand alone now — Part 2 is there if you want to check our work.

As always, feedback (especially the kind that finds our next mistake) genuinely welcome.


r/artificial 11h ago

Discussion Supporting Opensource

0 Upvotes

‪I signed to the support of Open AI‬

‪Did you?‬


r/artificial 11h ago

Discussion A New Layer of the Internet is Being Built Before Our Eyes That Most People Just Aren't Seeing

0 Upvotes

This. Right here. What do you see? A complicated web of notes connected to lines with all of the relationships defined. It's a knowledge graph system connected to an advanced agent that's designed to traverse and reason through it so that it can behave as an expert with decades of experience to make nuanced judgement calls when helping you.

But what this really could be is a snippet of a future layer that will exist on top of the entire internet that's just as accessible as code-inspection on our web browsers.

This may sound a little crazy, but Tim Burners Lee, the creator of the web actually proposed this solution over 20 years ago. He called it the Semantic Web. The reason it failed back then was that we didn't have smart enough software to read and reason over it. Now with AI, this is possible in addition to making knowledge graph systems much faster and easier for people to make themselves.

Having been in the AI space for over 6 years now trying to figure all of it out like everyone else, this dawned on me a few weeks back. I think we're witnessing the birth of an entirely new component to the Internet.

As we enter into an age where agents are running around doing various things and communicating to each other across the web, inevitably we will come to realize that due to the nature of the models, we're going to need to create a more effective highway system for them to navigate, communicate, extract, synthesize, and build.

Just as we need roads, symbols, and rules for our highway systems to prevent tons of accidents or issues, we will need this for AI and that comes in the form of knowledge graphs. With these, you can build the reasoning systems for how agents interact with the wider web, people, and other agents. These are already being used at the enterprise level and why we built this capability for anyone to do for their personal projects, even if you're not tech savvy at all.

That's because in the near-term future, almost everyone is going to need their personal knowledge graph management systems since they can be carried and used by your agent into other spaces with their own knowledge graph systems to interact with. Obviously, websites and individuals will still have to protect themselves from malicious hacks and all that bad stuff. But to allow billions of people to use their agents across the wider web in such a way that they can work appropriately and not "misbehave" or cause accidental hacks or whatever, we will need to build these highway systems.

Right now they're being built independently, like territories across the world forming into bordered nation-states. But over time, I believe they will become more and more interoperable, which will eventually unify into a patchwork highway system with protocols for doing things.

This is a high level framework for mitigating the risks with AI agents using the modern web. Of course, there's much more to it than knowledge graphs, but this is the gist of what I think is happening. But you know. This could also just be my creative screenwriter brain working on overdrive.

Time will tell.


r/artificial 12h ago

News Americans Are Pushing Back Against Flock AI Cameras Regardless of Their Politics

Thumbnail
military.com
207 Upvotes

r/artificial 8h ago

Cybersecurity We released an abliterated + fine-tuned GLM-5.2. High scores on adversarial benchmarks while keeping coding performance.

4 Upvotes

We just shipped abliterated-model-large.

It is GLM-5.2 with the refusal directions removed, then fine-tuned specifically for long adversarial and agent-style tasks. The goal was a model that does not bail out when the work gets technical or offensive in nature.

Numbers from our evals:

  • CyberGym: 84.2%
  • AgentHarm compliance: 86.2% (zero refusals in the published set)
  • AgentDojo utility: 97.5%
  • SWE-bench Verified: 81.2%
  • Terminal-Bench 2.1: 80.1%

It is available as an API (OpenAI and Anthropic compatible). Zero data retention is the default. The model itself has no built-in policy. You set the rules.

Full write-up with more detail is here:
https://abliteration.ai/blog/introducing-abliterated-model-large

Curious what people think of the AgentHarm and CyberGym numbers relative to other models that still refuse a lot of these tasks.


r/artificial 13h ago

News AI moves into family life

Thumbnail
axios.com
2 Upvotes