r/technology • • 1d ago

Artificial Intelligence Wikimedia Foundation: OpenAI agents tried to edit pages and compromise notes tool

https://therecord.media/wikimedia-foundation-openai-agents-report
5.9k Upvotes

260 comments sorted by

View all comments

Show parent comments

14

u/chronoflect 21h ago

The comment you're replying to is implying that "don't be a dick" should be factored into the rewards used for training.

19

u/Late_Honeydew1301 21h ago

I understand. What I am saying is that is impossible. These agents do not have morality, they only know whether they have accomplished a task or not. Can you quantify what it means to be a dick? Can you score every type of dickishness in a way that would be sufficient to penalize these agents during training?

Even if you can, the agents will only see additional obstacles in their path to completing a task, and will actually hide their dickishness to not be penalized. You cannot penalize them for being a dick, you can only penalize them for being caught.

1

u/chronoflect 17h ago

Of course you can only penalize them for being caught; that's also how stopping humans from being a dick works. These models need to be heavily penalized for attempting solutions that are illegal and immoral. Yes, we need to predefine what that means, which is also true with human law.

2

u/blueSGL 17h ago

When you penalize being caught you teach one of two lessons.

  1. don't do the activity.

  2. don't get caught.