Online Speech Is Now An Existential Question For Tech
Content moderation rules used to be a question of taste. Now, they can determine a service’s prospects for survival.
Content moderation rules used to be a question of taste. Now, they can determine a service’s prospects for survival.
Every public communication platform you can name—from Facebook, Twitter and YouTube to Parler, Pinterest and Discord—is wrestling with the same two questions:
How do we make sure we’re not facilitating misinformation, violence, fraud or hate speech?
At the same time, how do we ensure we’re not censoring users?
The more they moderate content, the more criticism they experience from those who think they’re over-moderating. At the same time, any statement on a fresh round of moderation provokes some to point out objectionable content that remains. Like any question of editorial or legal judgment, the results are guaranteed to displease someone, somewhere—including Congress, which this week called the chief executives of Facebook, Google and Twitter to a hearing on March 25 to discuss misinformation on their platforms.
For many services, this has gone beyond a matter of user experience, or growth rates, or even ad revenue. It’s become an existential crisis. While dialling up moderation won’t solve all of a platform’s problems, a look at the current winners and losers suggests that not moderating enough is a recipe for extinction.
Facebook is currently wrestling with whether it will continue its ban of former president Donald Trump. Pew Research says 78% of Republicans opposed the ban, which has contributed to the view of many in Congress that Facebook’s censorship of conservative speech justifies breaking up the company—something a decade of privacy scandals couldn’t do.
Parler, a haven for right-wing users who feel alienated by mainstream social media, was taken down by its cloud service provider, Amazon Web Services, after some of its users live-streamed the riot at the U.S. Capitol on Jan. 6. Amazon cited Parler’s apparent inability to police content that incites violence. While Parler is back online with a new service provider, it’s unclear if it has the infrastructure to serve a large audience.
During the weeks Parler was offline, the company implemented algorithmic filtering for a few content types, including threats and incitement, says a company spokesman. The company also has an automatic filter for “trolling” that detects such content, but it’s up to users whether to turn it on or not. In addition, those who choose to troll on Parler are not penalized in Parler’s algorithms for doing so, “in the spirit of First Amendment,” says the company’s guidelines for enforcement of its content moderation policies. Parler recently fired its CEO, who said he experienced resistance to his vision for the service, including how it should be moderated.
Now, just about every site that hosts user-generated content is carefully weighing the costs and benefits of updating their content moderation systems, using a mix of human professionals, algorithms and users. Some are even building rules into their services to pre-empt the need for increasingly costly moderation.
The saga of gaming-focused messaging app Discord is instructive: In 2018, the service, which is aimed at children and young adults, was one of those used to plan the Charlottesville riots. A year later, the site was still taking what appeared to be a deliberately laissez-faire approach to content moderation.
By this January, however, spurred by reports of hate speech and lurking child predators, Discord had done a complete 180. It now has a team of machine-learning engineers building systems to scan the service for unacceptable uses, and has assigned 15% of its overall staff to trust and safety issues.
This newfound attention to content moderation helped keep Discord away from the controversy surrounding the Capitol riot, and caused it to briefly ban a chat group associated with WallStreetBets during the GameStop stock runup. Discord’s valuation doubled to $7 billion over roughly the same period, a validation that investors have confidence in its moderation strategy.
The challenge successful platforms face is moderating content “at scale,” across millions or billions of pieces of shared content.
Before any action can be taken, services must decide what should be taken down, an often slow and deliberative process.
Imagine, for example, that a grass-roots movement gains momentum in a country, and begins espousing extreme and potentially dangerous ideas on social media. While some language might be caught by algorithms immediately, a decision about whether discussion of a particular movement, like QAnon, should be banned completely, could take months on a service such as YouTube, says a Google spokesman.
One reason it can take so long is the global nature of these platforms. Google’s policy team might consult with experts in order to consider regional sensitivities before making a decision. After a policy decision is made, the platform has to train AI and write rules for human moderators to enforce it—then make sure both are carrying out the policies as intended, he adds.
While AI systems can be trained to catch individual pieces of problematic content, they’re often blind to the broader meaning of a body of posts, says Tracy Chou, founder of content-moderation startup Block Party and former tech lead at Pinterest.
Take the case of the “Stop the Steal” protest, which led to the deadly attack on the U.S. Capitol. Individual messages used to plan the attack, like “Let’s meet at location X,” would probably look innocent to a machine-learning system, says Ms Chou, but “the context is what’s key.” Facebook banned all content mentioning “Stop the Steal” after the riot.
Even after Facebook has identified a particular type of content as harmful, why does it seem constitutionally unable to keep it off its platform?
It’s the “prevalence problem.” On a truly gigantic service, even if only a tiny fraction of content is problematic, it can still reach millions of people. Facebook has started publishing a quarterly report on its community standards enforcement. During the last quarter of 2020, Facebook says users saw seven or eight pieces of hate speech out of every 10,000 views of content. That’s down from 10 or 11 pieces the previous quarter. The company said it will begin allowing third-party audits of these claims this year.
While Facebook has been leaning heavily on AI to moderate content, especially during the pandemic, it currently has about 15,000 human moderators. And since every new moderator comes with a fixed additional cost, the company has been seeking more efficient ways for its AI and existing humans to work together.
In the past, human moderators reviewed content flagged by machine learning algorithms in more or less chronological order. Content is now sorted by a number of factors, including how quickly it’s spreading on the site, says a Facebook spokesman. If the goal is to reduce the number of times people see harmful content, the most viral stuff should be top priority.
Companies that aren’t Facebook or Google often lack the resources to field their own teams of moderators and machine-learning engineers. They have to consider what’s within their budget, which includes outsourcing the technical parts of content moderation to companies such as San Francisco-based startup Spectrum Labs.
Through its cloud-based service, Spectrum Labs shares insights it gathers from any one of its clients with all of them—which include Pinterest and Riot Games, maker of League of Legends—in order to filter everything from bad words and human trafficking to hate speech and harassment, says CEO Justin Davis.
Mr Davis says Spectrum Labs doesn’t say what clients should and shouldn’t ban. Beyond illegal content, every company decides for itself what it deems acceptable, he adds.
Pinterest, for example, has a mission rooted in “inspiration,” and this helps it take a clear stance in prohibiting harmful or objectionable content that violates its policies and doesn’t fit its mission, says a company spokeswoman.
Services are also attempting to reduce the content-moderation load by reducing the incentives or opportunity for bad behaviour. Pinterest, for example, has from its earliest days minimized the size and significance of comments, says Ms Chou, the former Pinterest engineer, in part by putting them in a smaller typeface and making them harder to find. This made comments less appealing to trolls and spammers, she adds.
The dating app Bumble only allows women to reach out to men. Flipping the script of a typical dating app has arguably made Bumble more welcoming for women, says Mr Davis, of Spectrum Labs. Bumble has other features designed to pre-emptively reduce or eliminate harassment, says Chief Product Officer Miles Norris, including a “super block” feature that builds a comprehensive digital dossier on banned users. This means that if, for example, banned users attempt to create a new account with a fresh email address, they can be detected and blocked based on other identifying features.
Facebook CEO Mark Zuckerberg recently described Facebook as something between a newspaper and a telecommunications company. For it to continue being a global town square, it doesn’t have the luxury of narrowly defining the kinds of content and interactions it will allow. For its toughest content moderation decisions, it has created a higher power—a financially independent “oversight board” that includes a retired U.S. federal judge, a former prime minister of Denmark and a Nobel Peace Prize laureate.
In its first decision, the board overturned four of the five bans Facebook brought before it.
Facebook has said that it intends the decisions made by its “supreme court of content” to become part of how it makes everyday decisions about what to allow on the site. That is, even though the board will make only a handful of decisions a year, these rulings will also apply when the same content is shared in a similar way. Even with that mechanism in place, it’s hard to imagine the board can get to more than a tiny fraction of the types of situations content moderators and their AI assistants must decide every day.
But the oversight board might accomplish the goal of shifting the blame for Facebook’s most momentous moderation decisions. For example, if the board rules to reinstate the account of former President Trump, Facebook could deflect criticism of the decision by noting it was made independent of its own company politics.
Meanwhile, Parler is back up, but it’s still banned from the Apple and Google app stores. Without those essential routes to users—and without web services as reliable as its former provider, Amazon—it seems unlikely that Parler can grow anywhere close to the rate it otherwise might have. It’s not clear yet whether Parler’s new content filtering algorithms will satisfy Google and Apple. How the company balances its enhanced moderation with its stated mission of being a “viewpoint neutral” service will determine whether it grows to be a viable alternative to Twitter and Facebook or remains a shadow of what it could be with such moderation.
The Australian leather house has opened an immersive four-day pop-up in Manhattan, unveiling its Bloom Collection and redefining what a product launch can look like.
Following the successful launch of its Palais Collection, MAISON de SABRÉ has unveiled a new modular handbag system offering more than 720 styling combinations.
AI doesn’t rebel—people design, deploy and profit from it. The real danger lies in allowing tech companies to escape accountability while shaping regulations that protect their dominance.
A wave of corporate warnings and technical disclosures has flooded the media, with headlines worrying over “swarms” of rogue artificial-intelligence agents launching “unprecedented” cyberattacks, outsmarting their makers, and inching toward a terrifying autonomy. The most revealing part of this narrative isn’t what the software did. It’s who is telling the story—and why. When corporate leaders publicly insist that the systems they financed, engineered and deployed are suddenly beyond their power to contain, skepticism isn’t only healthy; it is essential.
For years, Silicon Valley has drawn scrutiny from civil society and global regulators over tangible harms such as youth mental health deterioration and systematic privacy violations. Today, industry figures seem to be trying to change that public image. Loudly blowing the whistle on their own systems—just as two of the leading companies were preparing for massive initial public offerings—lets AI executives position themselves as a new generation of leaders who have come to terms with their societal responsibilities. They seem to want us to believe that they no longer want to “move fast and break things” but will instead stand as vigilant guardians between humanity and a technological apocalypse.
There is one glaring problem: Software doesn’t rebel. A mathematical model possesses neither intent, malice nor the will to defy its creators, let alone extinguish our species. AI is a human artifact, engineered for profit.
When an agentic model in an evaluation sandbox connects to an unauthorized server or executes an exploit, it hasn’t staged a coup. It has tried to meet the human-defined objectives set out before it through a path its designers failed to constrain. It’s the digital equivalent of the King Midas myth, in which the king’s ill-defined wish turns even his food and drink into gold.
That powerful experimental models were able to discover novel vulnerabilities and breach external systems isn’t a sign of a dangerous superintelligence but of human error or negligence. There is no sentient actor lurking in the weights to be reasoned with, feared or pacified. There are only human software engineers, product managers and corporate boards deciding which guardrails are worth the latency cost and which permissions can be skipped in the race to market.
Policymakers and voters need to resist AI exceptionalism. In any other discipline—from civil engineering to pharmaceuticals—courts and regulators treat a system failure as evidence of bad product design and inadequate safety testing. If an aircraft crashes, we focus on finding the engineering defect, correcting it, and enforcing established liability standards for the damage created.
By leaning on an anthropomorphic narrative, Silicon Valley attempts to repackage its specific human choices that led to experimental, powerful models behaving unexpectedly during tests as an existential peril. Elevating the issue to a cosmic scale leaves the public paralyzed and takes ordinary product accountability off the table.
In the cutthroat race for venture capital and market dominance, building guardrails slows down deployment. Grandstanding about uncontrollable power costs nothing and generates billions of dollars in free publicity, justifying stock prices, all while cultivating an aura of technological capability not only to build the frontier but also ultimately to rein it in.
Governments need to recognize regulatory capture when it stares them in the face. Tech leaders’ strategy looks transparent: Alarm Washington and Brussels into creating a regime in which only trillion-dollar incumbents with fully staffed compliance and safety departments can legally operate. By sitting at the policymakers’ tables before anyone else, these companies can help draft rules digging an impassable moat protecting them from open-source developers and upstart competitors, domestic or international. The real danger is in further concentrating the tech industry into the hands of only a few companies with deep pockets.
Beijing and Washington have brushed off those tech leaders’ calls, albeit for very different reasons. Chinese state media dismissed them as part of the “Cold War playbook” and intended to preserve U.S. dominance. Xi Jinping argued for exactly the opposite at the Brics Summit on Sept. 12, calling on Brics countries to “strengthen cooperation in the field of AI, encourage open source, openness, collaboration and sharing, and break new grounds and scale new heights.” President Trump, steeped in a doctrine of unfettered capitalism and technological supremacy, called fears that AI could destroy humanity a “hoax.” Vice President JD Vance warned that AI companies “begging the government to regulate them” looked like a “Trojan Horse.”
Striving to pursue its “European way” on AI and assert regulatory leadership, Europe, by contrast, welcomed the call. European Union President Ursula von der Leyen made this clear at the State of the EU speech last Wednesday and announced that the EU will invite “the main frontier labs for a discussion on how we can support ongoing industry efforts to pace the frontier.”
Europe has been here before. In an effort to lead global regulation and react to fears borne from ChatGPT, Europe rushed its landmark AI Act into law in 2024. Already the world’s most restrictive rulebook, the framework quickly proved too broad and complex to enforce. Stalled by implementation delays and concerns about European competitiveness, the EU postponed the law’s full rollout, leaving regulations uncertain.
AI should be regulated—risks exist and should be taken seriously. But governments need to act based on available evidence and verified facts, not corporate PR panic, the views of industry insiders, or the desire for quick political wins. The greatest danger facing society isn’t that software will awaken and overthrow its human masters. It is that we will allow the creators of the software to abdicate human responsibility for the systems they choose to build and help them pull up the ladder to market access behind them.