Let me try to phrase that differently. I'm not trying to legitimize bad stunts. I'm suggesting giving them a legitimate way of getting that PR instead.
Researchers, hobbyists and porn sellers working on chatbots. The definition of success is the Turing test. To give themselves goals and brag about those goals in media, they enter Turing test contests and try to get a "passing" score relatively to arbitrary definitions of pass.
Some will inevitably get overenthusiastic about the achievements. It will go int the press and that's where overpraising wolf crying comes from.
I propose to fixe this by giving contestants a legitimate target to aim for. Something that naturally produces PR fodder in 256 characters. "New Chatbot achieves 38% on the Eugene-13 Turing Test." That could mean something real and honest.
Some variants of the test are complete BS. Some are not. They may not be an indication of consciousness (that's what the one we've got is for) but beating our previous best does represent a legitimate milestone.
A lot of them could be interesting. Some variants of the test could focus on tasks/scenarios. For example, an AI receptionist. Others (maybe Eugene-13) could be a handicap version to the real Turing test.
As long as the butchers get to validate their own meat this won't happen. Nobody in the run for PR will allow you to rain on their parade with a 'not good enough'. Essentially what happened here is that the grades were lowered until there was a pass. If you set up shop next door and you refuse to lower the grades until there is a pass they'll go elsewhere, and claim victory.
Researchers, hobbyists and porn sellers working on chatbots. The definition of success is the Turing test. To give themselves goals and brag about those goals in media, they enter Turing test contests and try to get a "passing" score relatively to arbitrary definitions of pass.
Some will inevitably get overenthusiastic about the achievements. It will go int the press and that's where overpraising wolf crying comes from.
I propose to fixe this by giving contestants a legitimate target to aim for. Something that naturally produces PR fodder in 256 characters. "New Chatbot achieves 38% on the Eugene-13 Turing Test." That could mean something real and honest.
Some variants of the test are complete BS. Some are not. They may not be an indication of consciousness (that's what the one we've got is for) but beating our previous best does represent a legitimate milestone.
A lot of them could be interesting. Some variants of the test could focus on tasks/scenarios. For example, an AI receptionist. Others (maybe Eugene-13) could be a handicap version to the real Turing test.