top of page

Algorithmic Content Moderation: The Benefits and Risks of Using Artificial Intelligence to Govern Online Speech

  • Writer: Erin Smith
    Erin Smith
  • Jul 20
  • 9 min read

Updated: Jul 26


Social media platforms have emerged as some of the most dominant channels through which people communicate, access news and information, and express their opinions. Every day billions of pieces of content such as posts, videos, comments, and images are uploaded to social media platforms like Facebook, YouTube, Instagram, TikTok, and X, formerly known as Twitter. To moderate this vast amount of content created each minute, technology companies have turned to artificial intelligence (AI) solutions to assist with identifying and removing harmful content. Algorithmic content moderation refers to automated systems which can flag content that potentially violates a platform's policies before any human moderators have the opportunity to review the content. While these technologies have allowed platforms to respond more rapidly to content related to hate speech, terrorism, child exploitation and copyright infringement, they have also raised significant ethical, legal and political implications.


Algorithmic Content Moderation: Technical and Political Challenges in the Automation of Platform Governance , by Robert Gorwa, Reuben Binns, and Christian Katzenbach examines current trends toward reliance on automated moderation systems. They analyze both how these technologies work and their wider societal impacts. The article doesn’t suggest that content moderation should never be attempted with AI. It recognizes that automation is required to meet the vast scale of today’s social media sites. But the authors do argue that “algorithmic content moderation often shifts rather than resolves technical and political challenges.” (Binns et. al. 2018) These issues include transparency, fairness, accountability, and freedom of expression.


Why I Chose This Article

I picked this article because artificial intelligence is quickly becoming one of the most pervasive technologies present in our society. With that technology being focused heavily on digital communication and social media, I wanted to see how these areas overlapped. My educational background consists of studying communications, journalism, digital media marketing, and artificial intelligence. I've always been fascinated with how technology plays such an important role with mass communication. Every social media platform out there uses automated moderation.


Algorithms are deciding what people are allowed to see, share, and discuss online. It is also extremely topical because discussions about misinformation, political discourse, hate speech and online safety are more relevant now than ever. National governments are demanding technology companies do more to take-down harmful content rapidly. Users of these platforms are expecting them to allow free speech. These demands create an almost impossible task AI can not do on its own.


I picked this article as well because it did not just tout the technology nor bash it. It allowed the reader to understand how algorithmic moderation worked and the ethics behind why it may or may not work. The author's showed how there is no replacement for hard human decision making when it came to what speech stays online.


Thesis Statement

Algorithmic content moderation is necessary as humans cannot manually moderate all of the content that gets posted on social media. However, algorithmic content moderation should not be a substitute for human judgement. Content moderation algorithms lack context, transparency, fairness, and accountability. Artificial intelligence should assist human moderation, but should not be dictating what speech is allowed online.


Understanding Algorithmic Content Moderation

One significant merit that Gorwa, Binns, and Katzenbach’s article has is breaking down technical ideas of algorithmic moderation to simpler terms. Readers who may not have studied computer science can benefit and understand the rest of the article once they understand what algorithmic moderation means to the authors. According to Gorwa et al. (2020), algorithmic moderation encompasses “automated systems that classify or identify UGC and subsequently make governance decisions (e.g. delete posts, disallow uploads, restrict account functionality, surface ‘flagged’ content for human moderators)” (p. 1). Moderation systems at present can either follow one of two methods.


The first technique is hash matching. Hash matching identifies known bad content, like terrorist recruitment videos, child exploitation materials, or copyrighted works by assigning a digital fingerprint. When users upload files to a site, their fingerprint is matched with ones already in the database. If there’s a match, the system can automatically block the uploaded material or flag it for review. Hash matching is easy to implement and only identifies material that’s already been confirmed to violate a site’s policies.


The second way is through machine learning classification. Instead of looking for specific keywords/phrases, machine learning can be used to scan text, images, audio or video files to detect patterns and predict if newly posted content can be classified as hate speech, spam, harassment, terrorism or fake news. It does this by learning from tens of thousands or even millions of past examples that have already been reviewed and labeled by human moderation teams. The algorithm then tries to identify commonalities in future posts.


This is significant because there is a difference between these two methods. Hash matching is generally more successful because you are comparing things to previously confirmed pieces. Machine learning forces the algorithm to understand language nuances, context, sarcasm, comedy, culture, etc. which is something that even cutting-edge AI struggles with.


Why Platforms Depend on Automation

The writer argues that Big Tech firms don't have much option but to automate moderation. Given the sheer scale of user-generated content uploaded daily, human moderators can't keep up. Billions of daily posts cannot be screened in real time by humans. AI enables platforms to flag toxic content at scale, far quicker than would be possible by human moderators alone.


A real-world example outlined in the article took place following the 2019 terrorist attack in Christchurch, New Zealand. According to Facebook, 1.2 million copies of the attack video were blocked through automated hash-matching technology before they could even be uploaded. This showcases how Artificial Intelligence systems are able to help minimize the distribution of violent extremist media when these events happen so quickly. They also note how governments around the world have leaned on tech companies to accelerate takedowns of unlawful content by passing laws and regulations. Platforms have in turn grown reliant on algorithmic moderation in order to keep up, as the human-review alone isn't fast enough to comply with ever-stricter legal demands.


Not only legal coercion, but money talks too. Advertisers don't want their ads to be placed next to violent, hateful, or offensive content. Therefore, companies spend billions on automated moderation software to keep their advertisers happy, and their reputation intact.


Evidence Supporting the Authors' Argument

My favorite piece of evidence from Gorwa, Binns, and Katzenbach’s article is their argument that artificial intelligence was needed for content moderation because companies couldn’t hire enough humans to moderate every social media platform. With billions of users posting tens of billions of pieces of content daily, companies need algorithmic moderation to flag harmful content within seconds of being posted to minimize the amount of harmful content seen by millions of users.


The authors note that automated moderation works best for detecting content that has already been confirmed as violating. Platforms do not have to moderate each upload in real-time: Uploads can be scanned and crosschecked against databases of existing terrorist materials, child pornography, or copyrighted material, allowing companies to react nearly immediately and blocking repeat offenses.


Copyright Protection

Algorithmic moderation has been relatively successful in detecting copyright violations. The article covers YouTube's Content ID program, an early and advanced example of automated moderation. Copyright owners provide reference files of original audio or video to the system. YouTube can then automatically scan uploads to detect matches with the submitted files. Copyright owners can choose to block matching videos, monetize them on behalf of the uploader, or monitor their statistics rather than outright deletion.


Without Content ID, managing rights for millions of songs, movies and TV clips would mean forcing human beings to review each instance one-by-one. With several hundred hours of video being uploaded to YouTube every minute by its users, this task is simply impossible. As a result, automated technologies such as Content ID have become critical tools for intellectual property protection.


External research has found this to be true as well. The Organisation for Economic Cooperation and Development released a study that stated that technological platforms have moved towards AI operated programs due to the fact they operate faster, cut response time, and help facilitate an amount of information human employees wouldn't be able to work with on their own. Though the OECD admits there are ethical issues with AI, it has become necessary for automating tasks within massive digital spaces.


Detecting Hate Speech and Harassment

Another topic of conversation mentioned in the article is using machine learning to detect hate speech, harassment, bullying and abusive language. Instead of keyword lists, newer moderation systems learn patterns in language to determine how likely text or speech is to violate a service's rules. Facebook, Instagram, Twitter, and Google's Perspective API use machine learning models trained on enormous datasets of previously reviewed comments to forecast if posts will probably be offensive.


Scholars from outside this article agree that AI is increasingly helpful in identifying toxic speech. Schmidt and Wiegand (2017) state machine learning is more accurate at identifying hateful online speech than basic keyword screening as it takes larger sections of language into account. Machine learning systems are far from flawless, but they are effective and better than previous methods of moderation. Likewise, Fortuna and Nunes (2018) determined that natural language processing advancements have greatly benefitted automated hate speech detection. They reviewed that artificial intelligence systems are steadily improving as more training data are collected and language models improve.


This backs up the authors' argument that algorithmic moderation is playing an essential role in platform governance today. Automated systems might not catch everything, but they vastly accelerate the process of flagging toxic content.


Evidence Against the Authors' Argument

While Gorwa, Binns and Katzenbach rightly point out some key dangers around algorithmic moderation, there are equally valid counterarguments that A.I. will get better and fix some of today's problems. Advocates of A.I. moderation could point to how machine learning technologies get incrementally better as they're fed more training data, more nuanced language models, and more raw compute power to learn from. They would argue that we shouldn't view weaknesses of these systems today as necessarily permanent.


Another argument against this is that human moderation isn’t perfect. Humans get tired, burnt out, make inconsistent judgments, and are biased. Studies have found numerous content moderators leave their jobs with psychological trauma from constantly being exposed to violent, graphic, and otherwise degrading content (Roberts, 2019). By implementing AI, there would be less disturbing content for humans to review which would create a safer work environment for them and allow them to only handle the worst cases.


Fans of AI moderation also argue that automated systems are far more consistent than humans. Two people might enforce a platform's policies in very different ways based on their life experiences, culture or language fluency, or the time of day they're reviewing content. An algorithm will approach every decision the same way no matter how many millions of pieces of content it reviews. Consistency isn't fair consistency, but it can cut down on human error.


Another way AI is advancing is through becoming increasingly context-aware. Improvements to large language models and natural language processing allow computers to better understand sentence structure, emotion, sarcasm, and even context within conversations. Some scientists (Brown et al. (2020)) believe that current language models have shown the capacity for reasoning that was not thought possible years ago. In the future, moderation tools may one day be able to more accurately make contextual judgements than their predecessors. Others have pointed out that if moderation were only done by humans, abusive content could get far ahead of moderators than it already has. If there were ever to be a breaking news event, terror attack, or coordinated manipulation effort, moderation could only respond within minutes. In that time, harmful content can be seen by millions of people. Bots may be bad, but they're better than no bots at all.


My Analysis

Having read this article as well as further studies on the topic, I feel the authors make the better argument overall. While AI can be useful in managing massive amounts of content from the internet at rapid speeds, and it's unlikely our current platforms could run without some level of automation, the data shows that AI should be used to assist our decisions rather than make them for us.


The best use case for algorithmic moderation is detecting already-known bad content. Examples might include child pornography, terrorist recruitment materials, or copyrighted works. In these cases, automated moderation can vastly increase the speed with which platforms can respond. But it starts to fall apart when it comes to subjective conversation. Political debates, sarcasm, comedy, cultural gaps, art, and language evolution are tricky even for AI we have today.


Instead of trying to create independent moderation, platforms can build systems where AI flags content and trained employees step in when needed for edge cases and content of high public interest. Platforms should also be more transparent by explaining moderation actions, allowing outside audits, and building better appeal mechanisms.


Conclusion

Automated content moderation is one of the most pressing issues facing online communication today. With billions of users creating content online daily, it is inevitable that AI will continue to play a role in platform moderation. Gorwa, Binns, and Katzenbach effectively highlight how the benefits of automation, such as speed and scalability, are doubled-edged by risks to transparency, fairness, accountability, and democractic oversight.


The piece concludes by reminding us that technology will never take away challenging moral judgments. AI can help flag potentially offensive content, but it can't decide what society values or filter through disputes about freedom of speech. Those calls are still very human.


Looking forward, policymakers, tech companies, researchers and the general public should collaborate to ensure automated moderation promotes online safety and free speech. Transparency, continued oversight, and treating content moderation as both a technical and social issue will be key to accomplishing this balance.

References (APA 7th Edition)

Barocas, S., & Selbst, A. D. (2016). Big data's disparate impact. California Law Review, 104(3), 671–

732.

Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few-shot learners. Advances in

Neural Information Processing Systems, 33, 1877–1901.

Fortuna, P., & Nunes, S. (2018). A survey on automatic detection of hate speech in text. ACM

Computing Surveys, 51(4), 1–30.

Gorwa, R., Binns, R., & Katzenbach, C. (2020). Algorithmic content moderation: Technical and

political challenges in the automation of platform governance. Big Data & Society, 7(1), 1–15. https://doi.org/10.1177/2053951719897945

Roberts, S. T. (2019). Behind the screen: Content moderation in the shadows of social media. Yale

University Press.

Schmidt, A., & Wiegand, M. (2017). A survey on hate speech detection using natural language

processing. Proceedings of the Fifth International Workshop on Natural Language Processing

for Social Media, 1–10.



 
 
 

Comments


bottom of page