Highlights

Publications

Title page of the Medex paper

A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design (2025)

Haydn Thomas Jones, Natalie Maus, Josh Magnus Ludan, Maggie Ziyu Huan, Jiaming Liang, Marcelo Der Torossian Torres, Jiatao Liang, Zachary Ives, Yoseph Barash, Cesar de la Fuente-Nunez, Jacob R. Gardner, Mark Yatskar

Paper accepted into NeurIPS 2025

Figure 1 from the RAID paper: detectors succeed on default LLaMA output but fail with sampling and a repetition penalty

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors (2024)

Liam Dugan, Alyssa Hwang, Filip Trhlik, Josh Magnus Ludan, Andrew Zhu, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch

Paper accepted into ACL 2024

Diagram comparing a black-box fine-tuned LM with an interpretable concept-bottleneck prediction

Interpretable-by-Design Text Classification with Iteratively Generated Concept Bottleneck (2023)

Josh Magnus Ludan, Qing Lyu, Yue Yang, Liam Dugan, Mark Yatskar, Chris Callison-Burch

Explanation-based Finetuning Makes Models More Robust to Spurious Cues (2023)

Josh Magnus Ludan, Yixuan Meng, Tai Nguyen, Saurabh Shah, Qing Lyu, Marianna Apidianaki, Chris Callison-Burch

Paper accepted into ACL 2023. Video presentation on the left.

Random Projects

ReportCIS 522 Final Project (PDF)

Benchmarking moral decision making with various NLP architectures

Final project for CIS 522 – Deep Learning for Data Science

Evaluated a wide range of deep learning architectures ranging from CNNs to GPT3 in their ability to make complex moral judgements using a corpus of Reddit r/AITA posts

ReportComm 459 Final Report (PDF)

Modeling social contagion spread on Reddit

Final project for Comm 4590 – Social Networks and the Spread of Behavior

Completed class research paper analyzing how social contagions would spread in Reddit communities under different measures of graph centrality. This involved scraping and processing Reddit data in the range of hundreds of gigabytes and working with Prof. Damon Centola to use novel & more empirically accurate mathematical models of contagion.

ReportESE 3600 Project Report (PDF)

D&D Dice Reader Model

Final project for ESE 3600 – Machine Learning on Embedded Edge Devices

Created a model to read off the dice values in an image containing multiple Dungeons and Dragons dice. The pipeline involved combining multiple image recognition models (YoloNetV3 & MobilenetV2) and shrinking them efficiently enough to run on edge devices.

ReportNETS 212 Project Report (PDF)

Mini Facebook

Final project for NETS 212 – Scalable and Cloud Computing

Worked with a team of 3 to build an IaaS cloud hosted web app in the style of Facebook. App supports users with interests, friends, posts, walls, comments, live chat, and a personalized news feed based on user activity. Took on the role of project manager and designed & implemented the majority of the backend database API’s

ReportCIS 700 Final Project (PDF)

Interactive r/AITA Game using GPT-3

Final project for CIS 7000 – Interactive Fiction and Text Generation

Utilized GPT-3 to create an interactive game involving synthesized stories in the style of Reddit r/AITA posts

ReportCIS 545 Final Project (PDF)

Daily Pennsylvanian Topic Modeling

Final project for CIS 545 – Big Data Analytics

Performed topic modeling on the entire corpus of Daily Pennsylvanian articles to identify trends in the topics being discussed.

ReportCIS 520 Project: GeoGuessr (PDF)

GeoGuessr Geolocation model

Final project for CIS 520 – Machine Learning

Worked with a partner to create models predicting the location of a given image from Google street view using nothing but the image pixels. The best model utilizing EfficientNet architecture significantly outperformed previously published SOTA model, raising accuracy from 25.9% to 54.2% in a dataset predicting the state of an image taken within the US.

ReportNETS 150 Final Project Report (PDF)

Ivy League Subreddit Depression Ranking

Final project for NETS 150 – Market and Social Systems on the Internet

Analyzed the sentiment for various ivy league colleges to identify which communities were the most negative relative to others