Graduate Research Assistant: High Throughput Computing and Protein Machine Learning

FabAID is a new NSF-funded project building services to accelerate data-intensive and AI-driven science for the national research community (announcement). Our group's protein machine learning research is one of its science drivers. We use real modeling problems, like predicting how mutations change protein function, as a proof of concept for AI infrastructure that generalizes and can be deployed broadly. This position spans building the pipelines and using them to train models and make predictions.

What you'll do

  • Implement high throughput computing pipelines that run large-scale protein simulations across national computing resources
  • Curate training data: organizing simulation output and reformatting public experimental datasets
  • Train, fine-tune, and run inference on protein machine learning models, including our METL model
  • Work closely with the Center for High Throughput Computing team: co-developing software, running components of larger workflows, handing off data

An example project is building an automated pipeline that runs many protein stability simulations across national computing resources, pretrains a model with them, and fine-tunes the model with experimental protein function data. Concretely, this could involve an HTCondor DAG for Rosetta simulations, training the existing METL model architecture, and fine-tuning on deep mutational scanning data from ProteinGym.

You will own this project: debugging your own problems, reading the documentation, code, and literature you need in order to make good decisions, and setting your own short- and long-term goals. We meet weekly and also check in asynchronously to stay coordinated and provide feedback on your results and direction. AI tools are useful for this position. The expectation is that you take responsibility for all of your work and can defend anything you produce. This means checking output against documentation, primary literature, and your own tests, especially in the areas where you are newest and least able to catch a confident mistake.

What you need

Coursework in machine learning at the advanced undergraduate or graduate level covering topics such as classification and regression methodology, training and validation data splitting, choosing appropriate evaluation metrics, hyperparameter tuning, overfitting, and neural networks. Self-study can substitute for coursework if you can point to concrete evidence of equivalent coverage.

Everything else is teachable on the job: HTCondor and workflows, containers, PyTorch, and protein biology. Prior experience with any of them is beneficial, but none are strictly required. Please apply if you meet the machine learning requirement above, even if you have none of the rest.

This is probably not the right position if

  • You are primarily seeking first-author publications. We release all code, datasets, and workflows publicly, and you would get credit for the work you lead. Talks, posters, and press releases about the project are all possible. Publications may emerge, but they are not the goal of this work and are not guaranteed.
  • You want to train models without the data and infrastructure work. Pipelines and data curation are most of this job.
  • You need a fully remote arrangement.
  • You are a PhD student looking for dissertation research.
  • You are an undergraduate.

Appointment

50% research assistantship subject to university policies. The appointment starts at the beginning of the fall semester and runs through approximately June 2027, renewable contingent on funding and performance. Open to any current or admitted graduate student at UW-Madison.

In-person attendance is required at least 2-3 days per week and mandatory for individual and team meetings. The rest of your schedule is flexible.

How to apply

Email Professor Anthony Gitter with the subject line:

FabAID RA application: Firstname Lastname

In the body of the email, answer these three questions. 200 words maximum each. Please use a different example for each. Links to external resources strengthen your responses.

  • Tell me about something technical you built, solved, or figured out that you are especially proud of: code, a pipeline, an analysis, a dataset. What was your part in it, and why does it stand out to you?
  • Describe a computational problem that took you more than a day to debug. What was wrong, and how did you find it?
  • Describe a dataset you had to clean or reformat before you could use it. What was wrong with it, and what did you have to decide?

Also include in the body:

  • Your 3-5 most relevant courses: course number, institution, and one phrase on what it covered.
  • Your expected graduation date.

Attach your CV.

Review begins August 6 and the position is open until filled.

Equal opportunity and accommodations

UW-Madison is an equal opportunity employer. See the Equal Employment Opportunity Policy Statement.

If you need an accommodation for any part of this process, including the application or the interview, email Professor Gitter. You may also contact the Divisional Disability Representative for the School of Medicine and Public Health directly. More information for applicants is available here.

Drafted with Claude Opus 5.