Over the past decade, data has had an immense impact on the world. It has overhauled the technology industry, made immense shifts in the business and medical sectors, and even altered the way government functions. The media industry is no exception. This paper hopes to explore how the immense growth of data has affected reporters’ legal risks and legal rights. More specifically, it will examine how data has changed the legal risks for reporters’ newsgathering. For example, in the past decade there has been an increased risk of prosecution of journalists under the Computer Fraud and Abuse Act. The paper will also investigate how the data society has affected leaks and leak reporting. It will examine how the data boom has swelled the Freedom of Information Act (FOIA) process—and instigated the government to monetize its own data—and make journalists pay for information that should be free under the federal statute. Lastly, it will delve into the entangled relationship between robotics, AI, data, and the newsroom. In essence, this project hopes to understand how journalism responds to the growing changes in the information society.
Project lead: Victoria Baranetsky
Americans’ trust and confidence in the mass media is at an all-time low. Compounding this effect is the increase in “fake news.” It is becoming increasingly difficult for media consumers to distinguish between truthful and fictional news stories; readers of fake news reports often believe them, while readers of accurate reports question and mistrust them. A free press is a critical component of democracy, yet it seems to be in danger in the current age of media mistrust. To combat this trend, there have been recent efforts in the Natural Language Processing (NLP) community to use machine learning to automatically distinguish between “real” and “fake” news. This work is very important and will hopefully equip media consumers with the necessary tools to navigate the murky world of truth in media.
This project aims to study a complementary problem to fake news detection: trusted news detection. Instead of focusing on determining what is true or fictional, this study aims to discover the characteristics of trusted or believed text, regardless of the veracity of the text in question. Trust in media has been previously studied qualitatively: this proposed research is to our knowledge the first effort to quantitatively study trust in media on a large scale, using automated crowdsourcing, machine learning, and natural language processing methods. Further, this project proposes to analyze group-specific indicators of trust, to discover whether perception of trustworthiness varies across different categories of media consumers.
Project leads: Julia Hirschberg, Sarah Ita
Advances in artificial intelligence (AI) are influencing both the news industry and individual news consumption behaviors. The ability to convert structured data into a captivating story, indistinguishable from human-authored content, has large implications on the genesis and dynamics of audience segmentation. This project argues that audience fragmentation—accelerated by artificial intelligence—will be qualitatively different from that driven by either the multiplication of channels or on-demand personalized news consumption.
The purpose of the project is threefold: (1) segment today’s news audiences based on their current awareness, understanding and attitudes toward revolutionary changes that are being made in news production and distribution driven by artificial intelligence; (2) examine the audiences’ engagement with news and content powered by AI and automated journalism based on their current uses and gratifications; (3) identify the potentials and limits of AI-powered news and content to provide recommendations for ethically and efficiently incorporating AI technologies and requirements into the news ecosystem in a manner that best serves journalism and its audiences.
With these goals in mind, the current research examines how news audiences are segmented based on the beliefs held, the behaviors enacted, and the constraints faced concerning changes that are being made in news production and distribution powered by artificial intelligence and/or automated journalism.
This project will conduct two rounds of an online survey with adults in the United States. To segment news audiences, the survey data will be analyzed using latent cluster analysis (LCA), a statistical method for identifying unobserved subgroups within populations based on observed indicators. Unlike typical audience segmentation that is ad hoc and crude, the social scientific approach seeks to identify predictable groups based on the empirical observations that appear to be similar across a number of variables and subsequently develop an understanding of the underlying structure in terms of characteristics.
Project lead: Joon Soo Lim
In total this project produced three book chapters, two conference papers, an academic journal article, and a magazine article. The authors also wrote a total of 16 articles for various forms of news media including blogs as well as outlets like the Washington Post, Slate, and MIT Technology Review. Seven workshops were taught relating to the topics of the subawards including workshops related to investigating algorithms and designing news bots. Twenty-seven public presentations were given as part of the subaward, including keynotes to investigative journalists in Europe and Canada, as well as numerous panels and talks discussing algorithmic accountability and transparency. Five open source repositories were published.
The authors built and launched algorithmtips.org, which is a database of leads, as well as a community resource for journalists interested in beginning to investigate algorithms. The site attracts almost 1000 visitors per month (with high site engagement as people are looking through the database), and has attracted a handful of dedicated volunteers who are helping to expand the coverage of the database. The authors published a research paper at the Computation + Journalism Symposium in 2017 detailing how they built the database and laying out their plans for scaling up the efforts.
Diakopoulos' work relating to the ethics of algorithms was influential in crafting the Association for Computing Machinery's ethics guide relating to algorithmic transparency and accountability, which has since been adopted in Europe as well.
Project lead: Nicholas Diakopoulos
A new web dashboard that allows journalists to analyze, visualize, and interact with contractor data from governments.
Project lead: Alexandre Goncalves
A partnership between faculty and students in the Departments of History, Statistics and Computer Science at Columbia University, this project examined official secrecy by applying natural language processing software to archives of declassified documents to examine whether it is possible to predict the contents of redacted text, attribute authorship to anonymous documents, and model the geographic and temporal patterns of diplomatic communications.
Project leads: Nicholas Diakopoulos, Alexander Howard, Jonathan Stray
This report examines the operating biases of the new power brokers in society—algorithms—and the potential for accountability practices. Given the challenges to effectively employing transparency for algorithms—namely trade secrets, the consequences of manipulation, and the cognitive overhead of complexity—journalists might effectively engage with algorithms through a process of reverse engineering to understand the input-output relationships of an algorithm and develop stories about how that algorithm operates. This report proposes a practical method by which journalists can investigate algorithms, and develops best practices around exposing algorithmic bias and making algorithms transparent.
Project lead: Nicholas Diakopoulos
Team member: Jennifer Stark
This report brings together many fields to explore where data comes from, how to analyze it, and how to communicate results. Some of these ideas are thousands of years old, and some were developed only a decade ago—all of them have come together to create the 21st century practice of data journalism. This report serves as an introductory guidebook for journalists on the process of quantification, analysis, and communication using data.
Project lead: Jonathan Stray
This project aims to study the creation and consumption of automated news for forecasts of the 2016 U.S. presidential election.
Project lead: Andreas Graefe
Following four years of interviews with hundreds of editors, professors, reporters, technologists, government officials, and “hacker journalists,” this project details best practices of using data in journalism and offers solutions to remaining significant cultural, fiscal, and technical barriers to the adoption of data journalism and digital skills. The report suggests professional development along with statistical and scientific instruction for journalists, increased security practices around journalism, and a raised standard for accuracy and corrections. In addition, newsrooms must diversify their staff, and data journalists need to go beyond acquiring and cleaning data to understanding its provenance and source. Journalists must consider when it is appropriate to scrape data, access data, store it, or not—and understand that sensitive data will need to be protected with the same vigor that journalists have protected confidential sources.
Project lead: Alexander Howard
Team member: Jonathan Stray
This report evaluates the role of automated journalism, currently used to produce routine news stories for repetitive topics for which clean, accurate and structured data are available. Evidence available to date shows that while people rate automated news as slightly more credible than human-written news, they do not enjoy reading it since the writing is perceived as boring. While algorithms can describe what is happening, they cannot provide interpretations of why things are happening. Journalists are best advised to focus on tasks that algorithms cannot perform, such as in-depth analyses, interviews with key people, and investigative reporting.
Project lead: Andreas Graefe
This project centers on a prototype "build system" for data projects, or a beginner-friendly tool that enables incremental development of complex data projects and makes an entire data project reproducible from the source code alone.
Project leads: Mark Hansen, George King
Team members: Leopold Mebazaa, Gabe Stein
This project explores a new artificial intelligence tool—expert systems—that has the potential to aid journalists to quickly and efficiently discover public affairs stories in large public datasets. The logistics of performing data journalism have proved formidable for many news organizations. The project’s Creating a Story Discovery Engine would allow the local press to use it to write stories without having to fund its development or hire and manage a software staff.
Project lead: Meredith Broussard
