I Wouldn't Say Pangram is Broken, But I Would Say That It's Brittle

Recent testing shows that the AI detection tool Pangram easily produces contradictory results depending on text context, raising concerns over its reliability in professional settings.
I Wouldn't Say Pangram is Broken, But I Would Say That It's Brittle
it seems way too easy to provoke Pangram into contradicting its own results
So let me get to the nut of this thing before I do my usual meandering. Recently, someone accused me of using AI to write this old post from about a year ago, specifically highlighting the section about Ta-Nehisi Coates. I replied by saying that the post contained no AI writing; nothing I publish has been written or edited by an LLM. My accuser proceeded to say that the AI detection tool Pangram flagged that portion as 100% AI written with high confidence, a result I replicated. As the man who wrote that piece, I knew that I wasn’t guilty, but I also knew that wasn’t going to fly with other people. I was sufficiently annoyed that I paid for a Pangram license and found that, when I entered the entire essay, Pangram cleared it as 100% human written with high confidence. That is to say, the essay of about 5,000 words was declared 100% human written with high confidence even though it contains the section of about 300 words which was declared 100% AI written with high confidence. Subsequently, I’ve found that I can produce the same sort of inconsistency in the opposite direction - I can break the section that was called 100% AI into pieces that are then called 100% human written. So pieces that Pangram call 100% human written are part of a larger piece that Pangram calls 100% AI written which in turn is part of an essay that Pangram calls 100% human written, a Russian doll of contradictory results.
If nothing else, there’s a fundamental inconsistency here: it doesn’t make sense for a system to say that it has high confidence that an essay is 100% human written but also that a piece of that essay is 100% AI-written, again with high confidence, and also to say that the piece itself has subsections that are 100% human written…. These claims are mutually exclusive and must necessarily erode our confidence in the instrument. Experimenting, I’ve also found that I can pretty reliably get Pangram to declare large sections of text to be AI-written or human-written depending on the larger textual context I place those sections in - that is, I can take AI-generated text and embed it in human-generated text and have Pangram declare it 100% human, and I can take human-generated text and embed it in AI-generated text and have Pangram declare it 100% AI. I’ve also found that it’s pretty easy to induce false positives, that is, to write texts myself that Pangram identifies as AI-written. All of this makes me hesitant about the Pangram tool, and in particular the way many people use it, as a one-shot gotcha machine.
In a purely self-interested sense, the most obvious thing for me to do would be to point to this 100% human written outcome for the whole essay, say “See???,” and get indignant about having been accused. But I think there’s a lot to talk about here, and I think that people who write for a living have to have these conversations. These tools are being used right now, and with consequences, and we have to work this stuff out. So here goes.
The percentage meter seems clearly broken. One frustration of discussing Pangram results is that many people seem to think that the percentage that’s expressed is the confidence level - that is, that a 100% AI result is saying “we are 100% sure that this text is AI generated.” But that’s not what’s being measured there. The percentage is supposed to indicate what portion of the text is suspected to be LLM written. The degree of confidence is flagged there underneath “AI Generated,” although annoyingly the flag only appears when confidence is high, with no “Confidence Low” flag that ever appears, in my experience. This is all fine and good - the percentage AI generated and confidence numbers are different things that should be easy to interpret. The first problem is that many, many people are clearly interpreting “100% AI” to mean “100% confidence.” The second problem is that the percentage meter just does not appear to work! Anecdotally, a really suspiciously high portion of all Pangram outputs are 100%, whether 100% AI or 100% human. This is suspicious not only because a lot of LLM writing that gets published is likely hybrid, integrated with a writer’s own words, but also because the basic processes through which detectors like Pangram function seem likely to produce fewer polar outcomes. And a detector that spits out a lot of 100%s seems easier to break.
So let’s break the percentage meter. This is easy to do. Here’s a paragraph I’ve copy and pasted from a piece I wrote in 2017, which I hope you will accept as given was not written with the help of an LLM. I have appended three sentences to the end of the paragraph that were written by ChatGPT, using the original paragraph as a prompt. The human-written, pre-LLM section is 239 words, while ChatGPT’s portion is 71 words long.
So, ideally, Pangram would flag this as 23% AI generated; that is, indeed, what the exact percentages are by wordcount. Here’s what actually happens: That passage is three-quarters human written, dominantly human written, and yet Pangram says with high confidence that it’s 100% AI. I know there are a lot of people who would endorse a “one drop” rule with LLM assistance, but Pangram itself is saying that its technology operates a certain way, and it doesn’t appear to operate that way; certainly you can break it quite consistently.
And yet in dozens of trials I’ve never gotten a mixed report like this. They’ve all been either 100% human or 100% AI, even though I’ve been Frankensteining human and LLM writing together constantly for this exercise. For whatever reason, Pangram seems to want to find 100% results, even though users are clearly trusting it to detect what portion of a text is LLM generated. This all seems quite bad to me! This technology is being used in a way that has severe professional consequences for some people. It should work the way it’s supposed to work.
Source: Hacker News















