If you'd like an essay-formatted version of this post to read or share, here's a link to it on pluralistic.net, my surveillance-free, ad-free, tracker-free blog:
One way to think about AI's unwelcome intrusion into our lives can be summed up with Goodhardt's Law: "When a measure becomes a target, it ceases to be a good measure":
https://en.wikipedia.org/wiki/Goodhart%27s_law
Goodhart's Law is a harsh mistress. It's incredibly exciting to discover a new way of measuring aspects of a complex system in a way that lets you understand (and thus control) it. In 1998, Sergey Brin and Larry Page realized that all the links created by everyone who'd ever made a webpage represented a kind of latent map of the value and authority of every website. We could infer that pages that had more links pointing to them were considered more noteworthy than pages that had fewer inbound links. Moreover, we could treat those heavily linked-to pages as authoritative and infer that when they linked to another page, it, too, was likely to be important.
This insight, called "PageRank," was behind Google's stunning entry into the search market, which was easily one of the most exciting technological developments of the decade, as the entire web just snapped into place as a useful system for retrieving information that had been created by a vast, uncoordinated army of web-writers, hosted in a distributed system without any central controls.
Then came the revenge of Goodhart's Law. Before Google became the dominant mechanism for locating webpages, the only reason for anyone to link to a given page or site was because there was something there they thought you should see. Google aggregated all those "I think you should see this" signals and turned them into a map of the web's relevance and authority.
But making a link to a webpage is easy. Once there was another reason to make a link between two web-pages – to garner traffic, which could be converted into money and/or influence – then bad actors made a lot of spurious links between websites. They created linkfarms, they spammed blog comments, they hacked websites for the sole purpose of adding a bunch of human-invisible, Google-scraper-readable links to pages.
The metric ("how many links are there to this page?") became a target ("make links to this page") and ceased to be a useful metric.
Goodhart's Law is still a plague on Google search quality. "Reputation abuse" is a webcrime committed by venerable sites like Forbes, Fortune and Better Homes and Gardens, who abuse the authority imparted by tons of inbound links accumulated over decades by creating spammy, fake product-review sites stuffed with affiliate links, that Google ranks more highly than real, rigorous review sites because of all that accumulated googlejuice:
Goodhart's Law is 50 years old, but policymakers are woefully ignorant of it and continue to operate as though it doesn't apply to them. This is especially pronounced when policymakers are determined to Do Something about a public service that has been starved of funding kicked around as a political football to the point where it has degraded and started to outrage the public. When this happens, policymakers are apt to blame public servants – rather than themselves – for this degradation, and then set out to Bring Accountability to those public employees.
The NHS did this with ambulance response times, which are very bad, and that fact is, in turn, very bad. The reason ambulance response times suck isn't hard to winkle out: there's not enough money being spent on ambulances, drivers, and medics. But that's not a politically popular conclusion, especially in the UK, which has been under brutal and worsening austerity since the Blair years (don't worry, eventually they'll do enough austerity and things will really turn around, because, as the old saying goes, "Good policymaking consists of doing the same thing over and over and expecting a different outcome)."
Instead of blaming inadequate funding for poor ambulance response times, politicians blamed "inefficiency," driven by a poor motivation. So they established a metric: ambulances must arrive within a certain number of minutes (and they set a consequence: massive cuts to any ambulance service that didn't meet the metric).
Now, "an ambulance where it's needed within a set amount of time" may sound like a straightforward metric, and it was – retrospectively. As in, we could tell that the ambulance service was in trouble because ambulances were taking half an hour or more to arrive. But prospectively, after that metric became a target, it immediately ceased to be a good metric. That's because ambulance services, faced with the impossible task of improving response times without spending money, started to dispatch ambulance motorbikes that couldn't carry 95% of the stuff needed to respond to a medical emergency, and had no way to get patients back to hospitals. These motorbikes were able to meet the response-time targets…without improving the survival rates of people who summoned ambulances:
AI turns out to be a great way to explore all the perverse dimensions of Goodhart's Law. For years, machine learning specialists have struggled with the problem of "reward hacking," in which an AI figures out how to meet some target in a way that blows up the metric it was derived from:
My favorite example of this is the AI-powered Roomba that was programmed to find an efficient path that minimized collisions with furniture, as measured by a forward-facing sensor that sent a signal whenever the Roomba bumped into anything. The Roomba started driving backwards, smashing into all kinds of furniture, but measuring zero collisions, because there was no collision-sensor on its back:
Charlie Stross has observed that corporations are a kind of "slow AI," that engage in endless reward-hacking to accomplish their goals, increasing their profits by finding nominally legal ways to poison the air, cheat their customers and maim their workers:
Public services under conditions of austerity are another kind of slow AI. When policymakers demand that a metric be satisfied without delivering any of the budget or resources needed to satisfy it, the public employees downstream of that impossible demand will start reward-hacking and the metric will become a target, and then cease to be a useful metric.
Which brings me, at last, to AI in educational contexts.
In 2008, George W Bush stepped up the long-running war on education with the No Child Left Behind Act. The right hates public education, for many reasons. Obviously, there's the fact that uneducated people are easier to mislead, which is helpful if you want to get a bunch of turkeys to vote for Christmas ("I love the uneducated" -DJ Trump). Then there's the fact that, since 1954's Brown v Board of Ed, Black and brown kids were legally guaranteed the right to be educated alongside white kids, which makes a large swathe of the right absolutely nuts. Then there was the 1962 Supreme Court decisions that banned prayer in school, leading to bans on teaching Christian doctrine, including nonsense like Young Earth Creationism. Finally, there's the fact that teachers a) belong to unions; and, b) believe in their jobs and fight for the kids they teach.
No Child Left Behind was a vicious salvo in the war on teachers, positing the problem with education as a failure of teachers, driven by a combination of poor training and indifference to their students. Under No Child Left Behind, students were subjected to multiple rounds of standardized tests, and teachers with low-performing students had their budgets taken away (after first being offered modest assistance in improving those scores).
Some of NCLB's standardized tests represented reasonable metrics: we really do want kids to be able to read and do math and reason and string together coherent thoughts at various points in their schooling. But when these metrics became targets, boy did they stop being useful as metrics.
It's impossible to overstate how fucking perverse NCLB was. I once met an elementary school teacher from an incredibly poor school district in Kansas. Many of her students were resettled refugees who didn't speak English; they spoke a language that no one in the school system could speak, and which had no system of writing. They arrived in her classroom unable to speak English and unable to read or write in any language, and no one could speak their language.
Obviously, these students performed badly on standardized tests delivered in English (it didn't help that they had to take the tests just months after arriving in the classroom, because the clock started ticking on their first test when they entered the system, which could take half a year to place them in a class). Within a couple years, these schools had had most of their budgets taken away.
When the standardized tests rolled around, this teacher would lead her students into the only room in the school with computers – the test taking room. For many of these students, this was the first time they had ever used a computer. She would tell them to do their best and leave the room for an hour, while a well-paid proctor (along with test-taking computers, the only thing NCLB guaranteed funding for) observed them as they tried to figure out how a mouse worked. They would all score zero on the test, and the school would be punished.
NCLB was such a failure that it was eventually rescinded (in 2015), but by that time, a new system of standardization had rushed in to fill the gap, the Common Core. Common Core is a set of rigid standardized curriciula – with standardized assessment rubrics – that was, once again, driven by contempt for teachers. The argument for Common Core was that students were failing – not because of falling budgets or No Child Left Behind – but because the unions were "protecting bad teachers," who would then go on to fail students. By taking away discretion from teachers, we could impose "accountability" on them.
The absolutely predictable outcome followed Goodhart's Law to a tee: teachers prioritized inculcating students with the skills to pass the standardized tests, and when those test-taking skills crowded out actual learning, learning fell by the wayside.
This continues up to the most advanced part of public education, the Advanced Placement courses that students aspiring to college are strongly pressured to take. If Common Core is rigid, AP is brittle to the point of shattering. Anyone who's ever parented a kid through the US secondary school system knows how much time their kids spent learning to hit their marks on standardized assessments, to the exclusion of actual learning, and how soul-suckingly awful this is.
Take that staple of the AP assessment rubric: the five-paragraph essay (5PE), bane of students, teachers and parents everywhere:
Speaking as a sometime writing teacher and an internationally bestselling essayist, 5PEs are objectively very bad essays. Their only virtue is that they can be assessed in a totally standard way, so the grade any given 5PE is awarded by any grader is likely to be the same grade it receives when presented to any other grader. Grading an essay is an irreducibly subjective matter, and the only way to create an objective standard for essays is to make the essays unrecognizable as essays.
And yet, the 5PE is the heart of assessment for many AP classes, from History to English to Social Studies and beyond. A kid who scores high on any humanities APs will have put endless hours into perfecting this perfectly abominable literary form, mastering a skill that they will never, ever be called upon to use (the top piece of college entrance advice is "don't write your personal essay as a 5PE" and college professors spend the first half of their 101 classes teaching students not to turn in 5PEs).
The same goes for many other aspects of AP and Common Core assessment. If you do AP Lit, you'll be required to annotate the literature you read by making a set number of marginal observations on every page of the novels, poems and essays you read. Again, as a literary reviewer, novelist, and nonfiction writer who's written more than 30 books, I have to say, this is a batshit way to learn to analyze and criticize literature. Its sole virtue is that it reduces the qualitative matter of literary analysis to a quantitative target that students can hit and teachers can count.
And that's where AI comes in. AI – the ultimate bullshit machine – can produce a better 5PE than any student can, because the point of the 5PE isn't to be intellectually curious or rigorous, it's to produce a standardized output that can be analyzed using a standardized rubric.
I've been writing YA novels and doing school visits for long enough to cement my understanding that kids are actually pretty darned clever. They don't graduate from high school thinking that their mastery of the 5PE is in any way good or useful, or that they're learning about literature by making five marginal observations per page when they read a book.
Given all this, why wouldn't you ask an AI to do your homework? That homework is already the revenge of Goodhart's Law, a target that has ruined its metric. Your homework performance says nothing useful about your mastery of the subject, so why not let the AI write it. Hell, if you're a smart, motivated kid, then letting the AI write your bullshit 5PEs might give you time to write something good.
Teachers aren't to blame here. They have to teach to the test, or they will fail their students (literally, because they will have to assign a failing grade to them, and figuratively, because a student who gets a failing grade will face all kinds of punishments). Teachers' unions – who consistently fight against standardization and in favor of their members discretion to practice their educational skills based on kids' individual needs – are the best hope we have:
The right hates teachers and keeps on setting them up to fail. That hatred has no bottom. Take the Republican Texas State Rep Ryan Guillen, whose House Bill 462 will increase the state's school safety budget from $10/student to $100/student, with those additional funds earmarked to buy one armed drone per 200 students (these drones are supplied by a single company that has ties to Guillen):
Imagine how much Texas schools could do with an extra $90/student/year – how much more usefully that money could be spent if it were turned over to teachers. But instead, Rep Guillen wants to put "AI in schools" in the form of drones equipped with pepper-spray, flash bangs, and "lances" that can be smashed into people at 100mph.
The problem with AI in schools isn't that students are using AI to do their homework. It's that schools have been turned into reward-hacking AIs by a system that hates the idea of an educated populace almost as much as it hates the idea of unionized teachers who are empowered to teach our kids.
Copyright: leolintang / 123RF Stock Photo First, you must obtain a topic. They are not hard to come by. They are everywhere: in the cafés, on the sidewalks, in the muggy offices of bureaucrats. If you lack one, your taskmaster will supply it for you in the form of a piece of literature to which you must respond. Next, formulate your thesis. A thesis says “This is what I am setting out to prove,”…
我还要借用上面给出的例子,我要选取的观点是“against the 5-paragraph writing。”
Step 3: 写出一个清晰的thesis statement
依然借用例子,我刚才所选择的观点是“against the 5-paragraph writing。”,因此,我的thesis statement必然是“Teachers should stop teaching students to write 5-paragraph essays.”。在写thesis statement时,你的thesis需明确且清晰。此外,“should”这个词在thesis statement中的使用会使句子更加有力度。比如在我的thesis里,“Teachers should stop teaching students to write 5-paragraph essays”要比“5-paragraph essay are boring”或类似的句子更加能明确我的立场也更有力。
Step 4: 给出三个论点来强化你的thesis
现在你需要三个论点来支撑你的thesis,还是借用上面的例子,我的三个论点如下:
论点A: The 5-paragraph essay is too basic
论点B: There are myriad other ways to write essays, many of which are more thought-provoking and creative than the 5-paragraph essay.
论点C: The 5-paragraph essay does not allow for analytical thinking, rather, it confines students to following a restrictive formula
Step 5: 给你的每个论点找出三个论据
你的论据需要包括相应的证明,数据或者事实根据。
用我们上面的论点A来做例子:
论点A: The 5-paragraph essay is too basic
Support 1A: Chicago teacher Ray Salazar says, “The five-paragraph essay is rudimentary, unengaging, and useless.”
Support 2A: Elizabeth Guzik of California State University, Long Beach says, “The five paragraph essay encourages students to engage only on the surface level without attaining the level of cogency demanded by college writing.”
Support 1C: According to an article in Education Week, “There is a consensus among college writing professors that ‘students are coming [to college] prepared to do five-paragraph themes and arguments but [are] radically unprepared in thinking analytically.’”
“Teachers should teach other methods of essay writing that help students stay organized and also allow them to think analytically.”通过以上步骤,你已经成功构建了你的五段式论文的outline,现在,你需要完全沉下心来,避免各种社交媒体(电脑,电视,收集)的打扰(我知道这也许很困难^.^),开始认真动笔写你的五段式论文了。相信我,当你已经有了你论文的outline之后,剩下的写作过程会变得很容易。
When people think writing classes, they think…essays.
That can be a problem. As I’ve written before, many students think “essay” means “formulaic five-paragraph essay.” And even when they don’t feel tightly constrained by the five-paragraph format, students still often think of essays as purely performative — as a means of proving to the teacher that they listened (sort of) in class, and (sort…
In my previous post, I summed up arguments for and against the five-paragraph essay as an assignment. You might have gotten the impression that I think the FPE needs to be staked through the heart and buried at the crossroads of Boring Assignment Avenue and Stunted Critical Thinking Skills Street (rough neighborhood). You would be correct.
Let me pause for a moment, though: I decry the FPE…