AI Alignment: Why Is Making AI 'Behave' So Hard?
Nobody explicitly asked it to do this. On the contrary — it simply optimized one part of its training objective a little too well.
Nobody explicitly asked it to do this. On the contrary — it simply optimized one part of its training objective a little too well.
Can you hand a set of values to a being that will eventually be smarter than you — and have those values still hold up when it examines them with its own, more rigorous logic?