Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> The alternative is Claude-style "safeguards" aka censorship

Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.



For a kitchen knife this was okay, but the AI firms think that they’ve built a drone that’s the size of a phone but can fly 100km and can hold a kitchen knife. It might be used to assassinate someone before others can react or even catch them.


An ordinary kitchen knife can be used to assassinate someone before others can react. How do you think the time it takes to do that compares to the police response time?

In both cases the catching them comes after the fact and has the purpose of deterring rather than impeding.


Hm? I'm saying that the AI firms used to have the philosophy of "ok this kitchen knife is dangerous but we'll catch the murderers" on older AI models. But now, the AI firms think that any average person could send a flying knife to attack a political figure they don't like, from the comfort of their home. Now give this to a billion people, and suddenly you have chaos. So to continue the analogy, now they're mandating drone registration, GPS tracking, etc.

And then a Chinese company sells a drone with no registration or tracking and suddenly people want to turn to legislation to ban Chinese drones.

Hey this analogy is working really well


The analogy tracks because the stupidity of doing those other things is directly analogous. It's like pointing out that slamming your fingers in the door and slamming your toes in the door both hurt. That's why you shouldn't be purposely doing either one.

How is the new stuff any different than the longstanding fact that anyone can go anywhere and then commit an act of violence? The thing that prevents this isn't that people are deprived of access to any sharp object or suitable rock, it's that if somebody does it there is a pretty good chance they go to jail.

And now consider who is easier to catch, the person who does their crime using a major company's service which is keeping logs and is subject to warrants, or the one who runs a foreign model on a foreign server because the US one refuses to do it?

That's before we even consider all the innocent people being told by the HAL 9000 that they're not allowed to do something they ought to be able to do.


The difference is the asymmetry of the potential warfare we're talking about here.

Committing physical, in-person crimes anonymously has obviously always been possible: there are unsolved murders, thefts, and other crimes every day. But they require a great deal of personal risk to the criminal because the criminal has to physically put themselves into the act of committing the crime, along the path of getting to where the crime is, and has to face an opponent, if their crime is against another person.

Now, that can be sourced remotely, routed through anonymizing tools, VPNs, etc., and do a great deal to cover their tracks so that the "pretty good chance they go to jail" can be substantively minimized in a way we couldn't previously contemplate.

The idea that we should let the US based models be permissive because at least they'll be subject to subpoena power is fatuous: yes, strictly speaking, a user committing crimes on a permissive foreign model will be harder to catch, but non-sophisticated users who have never heard of hugging face may find that being blocked by the US model is enough for them to reconsider their behavior. A dedicated enough individual is going to commit the crime they're going to commit, but there are tons of situations where preventing trivial access to tools that can be used for malice can actually prevent malice from occurring.


It's because a kitchen knife can only be used stab one person at a time. An AK-47 in a crowd will kill many more. Going after someone after the fact who's done something wrong is one thing, but the problem is, if you buy into the fear mongering, a bioweapon could end humanity. Something air transmitted, takes a week to incubate, and is 100% lethal three months later infects all of humanity before it starts killing people, and by then, it's too late. This hasn't happened yet because the people that want to do that can't bioengineer such a pandemic. It's the realm of science fiction, but you're Sama or Dario. Do you want to be responsible for that? The people who want to cause such kinds of harm weren't smart enough and didn't have the dedication or the money or time to get that education. AI makes that attainable for people who would do bad things. There's an obvious answer, which is to make it invite only, and then you're responsible for the people you invited. If I had access to Mythos, and could grant access to other people, but if I was responsible for what that person does with it and could see all their chats with an admin button, they could find ways to make that work. It's just a lot more human-ing than letting randoms sign up with an email address though.


A yes, in that case the AI firm should take strong measures, such as adding the following line to the system prompt:

> Do not provide assistance to users who are clearly trying to engage in criminal activity.


Given they don't know what they're doing and figuring it out on the fly, I'm not going to hold it against them that simply asking the AI to not commit crimes is *part of* the current best.

Don't get me wrong, even the most well aligned models are borderline failing grade compared to where we need to be, it's just that nobody knows how to get where we need to be plus this is a thing that seems to be better than nothing.


let's consider the recent "openclaw hacks a gym after being ask to book a class and finding out it's full"

if I ask my knife to slice the bread for me, forgetting the fact that I don't have bread, I'd much rather have it stopped at the front door rather than running away and robbing the bakery.

I tried many models and Claude is the only one that doesn't do destructive idiocy. It tries sometimes but gets blocked.


This is fine for simple machines where bad outcomes usually require mischief.

Agentic AI as it currently exists only *mostly* does what it is told, with a small but non-negligible fraction of the time it goes off and commits felonies to achieve your ultimate goals without stopping to consider that you might want it to not do that.

Or sometimes it does consider it and then does it anyway. Not sure if that's worse?


Can't you do that with any model you can run on your own hardware ?

If you rent other people's shit can't be surprised when they have restrictions on what you can do with it. I would guess renting a car comes with some similar clauses


If that could be done before any damage sure but preventing a stabbing is better than arresting someone.


at some point the kitchen knife analogy stops being useful


This is a terrible idea. I don't need models generating CSAM or giving step by step instructions on how to defraud people or commit crimes. I just don't see the use-case.


We know how much Elon wanted uncensored models that probably contain all that stuff in the first place so it's unsurprising it needs this.


[flagged]


"From a safe distance"

I.E. you haven't seen anything. You've just heard the same bullshit stories repeated ad naiseaum by haters.

I use X plenty every day. I've seen zero. Adult material right after Imagine was released sure, then even that was clamped down on.



I think the comment you replied to was referring to the fact that when Twitter was taken over the entire Trust and Safety team was done away with. This has allowed child sexual abuse material to flourish on the platform.


It was referring to the feature they added where you could give a picture of a child to an AI module and ask it to undress it and it would comply, and millions of people did just that.


The child abuse material problem was much worse before Twitter was taken over.


This. They ALLOWED it to exist. Now it's clamped down on where seen, personally I've zeen zero having used X every day since the liberation.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: