TL;DR
- I built a data preprocessing agent with my teammates Max, Sean, and Shreya.
- We built a benchmark suite called KaggleBench. The novel idea is to evaluate agents on high-quality, human-expert data preprocessing workflows selected from Kaggle, rather than synthetic data.
- We built an agent with a workflow and tools tailored specifically to this task. We also built the agentic loop from scratch to understand the internals of AI agents.
- It was fun!
You can take a look at the details of the benchmark and the agent!
How it started
In CS 498 AI Agents in the Wild, we worked on a project where we designed both an AI agent and a benchmark from scratch.
We spent the most time choosing the topic. Since we were going to spend a semester on it, we wanted the result to be meaningful. Even though this was a student project funded out of pocket, we wanted the outcome to be high quality enough to be proud of.
So we made decisions using three rules.
- The topic should only be solvable with an agentic system. If a rule-based system or traditional machine learning could solve the problem well enough, we should avoid it.
- The topic should be measurable and quantifiable. We needed a dataset and a metric that could be justified reasonably. This matters because we wanted to show that our system works better than a simple baseline or prior work, and more importantly, that we can evaluate the system in a way that makes sense. If we cannot measure progress, we cannot improve the system in a disciplined way.1
- It should be possible to frame the topic as a problem with real industrial interest. The topic itself did not have to be industry-specific, but the core idea had to matter in industry. A research assistant agent maps naturally to information retrieval and summarization from user queries. A multi-turn customer service agent maps to chatting under policy constraints while retrieving information. These patterns are useful to many companies.
After comparing several ideas by pros, cons, feasibility, and viability, we decided to study whether an agent can solve data preprocessing tasks. The topic fit our rules. First, it is essentially a coding problem, so it is hard to solve without an agent. Second, by using Kaggle, we thought we could turn it into a quasi-quantifiable evaluation problem. Third, data science agents are already a visible industrial use case. OpenAI has written about its internal data agent, and Databricks also documents a data science agent workflow in notebooks.23
That is how we started working on the problem.
KaggleBench
For any system, defining quantitative evaluation first is important. Many LLM-based features or agent projects end with qualitative demos. The generated text looks impressive, and the project stops at subjective impressions. That is fun, but it does not create a path for improvement. Without a measurement setup, it is hard to know whether the next version is actually better.
At first glance, data preprocessing sounds like a small subset of a data science agent. But if you think about it carefully, it is not that different from a coding or math-solving problem. A data science project starts from ambiguous context, an imperfectly specified objective, and an open-ended search space. There is often no single perfect answer. Given how much money is being invested into coding agents, it would be unrealistic for a student team to solve the fully open-ended version in one semester.
So we had to scope the problem down. We reframed it as a multiple-choice decision problem. Given a Kaggle competition, its dataset, a set of candidate preprocessing actions, and a set of actions already taken, the agent has to decide which action should be added or removed.
This framing makes evaluation much easier because the problem is no longer fully open-ended. The hard part becomes generating good candidate actions. A good candidate action has to be clearly right or clearly wrong. It should not depend too much on subjective judgment.
For that, we used human-expert Kaggle notebooks as the source of truth. The details of how we built the dataset and evaluation suite are in the KaggleBench post.
Data preprocessing agent from scratch
One constraint from the class was that we could not use a well-known agent SDK such as LangGraph. That was a useful constraint. At a high level, an agent can sound simple. You call an LLM in a loop, parse tool calls, execute tools, and feed the result back into the next step.
The real difficulty is not the loop itself. The important parts are the harness around it and the durable execution behavior. The agent needs the right tools, the right prompt structure, the right state representation, and a way to keep making progress without losing track of the task.
There are already many emerging practices around agent harness design. We selected a few that fit our setting and built the architecture around them. We also implemented the agentic loop ourselves so that we could understand what is happening inside the system instead of treating an SDK as a black box.
We evaluated the agent against several baselines, including rule-based heuristics and a naive agentic loop. The baselines were designed so that our final method could be understood as a combination of smaller design choices. That made the comparison work like an ablation study. We also compared against Claude Code to see how our task-specific agent performed against a commercial general-purpose coding agent.
The detailed implementation and results are in the agent architecture post.
What I Learned
I got two main things from this project.
First, I learned a lot about agent architecture. Building an agent from scratch gave me a clearer big-picture understanding and more confidence. Of course, what we built was still an application-layer system that combined existing components. Each layer hides a lot of depth. Frontier labs know how to train LLMs with reinforcement learning for agentic use cases. LLM API providers do serious engineering around caching and request orchestration to use GPU resources efficiently. Agent harnesses for long-running tasks need many layers of reliability engineering. This project gave me a practical taste of what has to be considered at the application layer to build an end product.
Second, I got much better at using agents. Codex played a huge role in the development process. Because of it, I could spend more time on planning and decisions, which raised the quality of the final project. But it was not enough to simply hand off requests. I had to check persistently whether the thing I wanted had actually happened. That check could not only be conversational. I needed mechanical checks, tests, or direct inspection of outputs.
That process helped me feel the limitations of coding agents in a way that metrics alone cannot show. It was also interesting to experience model improvements during the project, from Codex 5.3 to 5.5.
Reference
-
Commonly attributed to Peter Drucker. ↩