
The Strategic Impact of Higher Education Accreditation Software
Accreditation is more than just a regulatory hurdle; it’s a critical stamp of approval that validates an institution’s credibility, academic integrity, and funding eligibility. But
Benchmarking sounds simple: compare your campus to others. But turning millions of public data points into a strategy campus leadership will actually use is a very different problem.
In this episode, host Debbie Phelps talks with Jon Gallegos, Data Analyst in the Institutional Research office at WSU Tech, about his 2026 AIR Forum poster presentation, "The Benchmarking Process from Initiation to Implementation." Jon walks through the exact technical and strategic steps his team used to build a benchmarking data store from the ground up — starting with leadership buy-in, standardizing on IPEDS survey data, and structuring longitudinal data in a Snowflake repository using Python.
The conversation goes deep on the technical side too: how clustering and cosine similarity can surface true peer and aspirational institutions beyond basic groupings, how Jon built a custom "Outcome Efficiency Index" connecting campus spending to student awards, and the Power BI design choices — intentional color, simplified matrix visuals — that make benchmarking data something colleagues actually use.
00:38Thank you for joining me today for another episode of Data Stakes, where I have conversations with professionals who work directly in the institutional research or effectiveness field, or are data-adjacent in their role in higher education. Today's conversation will focus on the importance of creating an integrated strategy so you can manage your benchmarking data. My guest today is Jon Gallegos. Jon is a data analyst in the Institutional Research office at Wichita State University Tech, a two-year technical college in south-central Kansas. His education and expertise in statistics, along with tools like SPSS, Python, Excel macros, and SQL, along with experience with AI models, help him give meaning to his campus's data. Today we're going to focus on his recent poster session at the 2026 AIR Forum, entitled "The Benchmarking Process from Initiation to Implementation." Jon, thanks for coming on Data Stakes.
01:50Thanks so much for having me — it's very exciting that you invited me, and this project has been in development for so long that it's nice to have a clear moment of, we've made it, it's here. That's a good feeling.
02:07That's always the news you want to share at the end of a session, especially at a conference — not "more to be continued," but "we made it, here's how you can do it." So let's dive in. We'll talk about benchmarking in general, then wade into the technical how-tos for how our listeners can learn to create a data store with structured data, select and group peers, and develop a report to communicate the results across campus. So — how does a new IR professional, or even a seasoned one tasked with benchmarking for the first time, start with purpose?
02:49That's one of the most important parts of this project, because it really guides what approaches you take later — what metrics you use, what peers you evaluate. I think that's something we could have done better, and maybe why it took a little longer than we wanted — the purpose just wasn't there in the beginning. One of the best places to start is with leadership. For us, this was a project specifically requested by the leadership team — they brought it to us and said, this is what we want out of it. That gave us purpose and guided what the project was meant for, and the buy-in was already there. That's part of the difficulty I heard from a lot of people at AIR Forum — the buy-in just wasn't there; they'd produce something and everyone would ask, why are we using this? So identifying what your leadership team is actually interested in is where I'd start — everyone's situation is going to be different, but that gave us a real sense of what this project was going to be.
04:01That's critical — having a plan, knowing where you're headed, maybe even focused on one or more data sets, and then going to leadership. Having the influencers on your campus behind you means it becomes an organizational strategy rather than a "nice add-on" that everyone can ignore. So you decided to make survey data from IPEDS — the Integrated Postsecondary Education Data System — your primary data set. Why?
04:45That was mostly Kristen's influence — she's the executive director of IR at WSU Tech, and she pointed me toward IPEDS. I have a better understanding now of why, but a few reasons stand out. First, it's all publicly available, which gives a lot of freedom — I don't have to worry about who has permission if we put it out campus-wide. Second, we have a lot of what I'd call "IPEDS reporters" — people in finance, people in financial aid who submit IPEDS data — and they already know the variables and metrics involved, which was key when going to leadership, because part of IR's job is saying, here's what we can provide. And third, it's all standardized: every college uses the same definition for first-year retention, graduation rates, and so on — which matters enormously when you want to compare against other two-year schools using the exact same definitions.
06:20That's very similar to the strategy I used at Cowley College — and listeners, Jon and I are located about an hour apart. IPEDS gives you consistent terminology, consistent metrics, and stakeholders on campus who already have the context you'll need when you share reports. I've always been a fan of it — it's free, it's public, and when I started my career back in 2007 at a smaller institution with very little budget, I needed something robust that wouldn't disappear in a year or two. So — how did you build your benchmarking data store? I saw your poster and I'm intrigued — give us the programming language, the data repository, the metric joins, all of it. The nerds and geeks are waiting.
07:49This is probably the piece I got to influence the most — a lot of the metrics came from staff who understand the nuances of higher ed better than I do, which is how it should be, but the technical build was mine. The first step was pulling IPEDS data into our own system using Python — it's all free, and you don't even need an API key; you just grab the URL and it sends back a zip file. You can also just go to the IPEDS website and download the files manually, and honestly I'm almost favoring that method now, because IPEDS tends to change code values for different frequencies and metrics — you have to go through each survey year manually and check whether the definitions still match. Financing variables changed between 2019 and 2020, for instance — some disappeared, some changed — so I had to make sure we matched the definition we were using for our composite financial index, which finance had already worked out. There's more than one way to get the data; I used Python, but you can absolutely pull the tables manually. Structurally, IPEDS collects by survey and by survey part — the fall enrollment survey, for example, has parts A, B, C, and D — so I'd union those tables together across years, matching variable names, which essentially builds one table five times as big if you're pulling five years. That gives you a longitudinal way to track time across your metrics. All of those tables go into Snowflake.
10:31So that's your data repository — a Snowflake data lake — which you then query to build your report. Other than the composite score, I saw something on your poster about a "benchmarking data score" — do you recall what that was? And while you're thinking, listeners, I want you to know we're going to share out Jon's documentation and an image of the poster he presented at AIR.
11:13This was a little self-serving too — my GitHub repository lets me track how many people view, clone, or branch off it, so I'm more than happy to share it. As for the "data score" — I'm honestly not sure what that's referring to; I do have something called the data store, which was one of my points, so that might be what we're looking at.
11:46That might just be a typo in the script I gave you — sorry about that. When I read through the rest of your presentation, I saw familiar calculations like first-year retention rate and graduation rate, but I also noticed you created one called the Outcome Efficiency Index. What is that, and what purpose does it serve?
12:10This is where leadership comes back into play — they specifically wanted to see a measure of how much we're spending per award, or more precisely, per award per student. The IPEDS variable I'm looking at is F1C0191, total expenses, and then CSTOTLT, total awards. We take that ratio and multiply it by a hundred thousand, so you get how many students receive awards for every hundred thousand dollars spent. It's very specific to us — another school might want to know how many students graduate per fifty thousand dollars, or per half a million — whatever makes sense for them. I've included the IPEDS variables themselves, though you'll have to look up the exact definitions.
13:35So you're blending two IPEDS surveys — the completion survey, which gives you the number of awards, and the finance survey. I love that, because one thing I noticed at AIR Forum this year is the data community talking more about financial metrics. Having worked in a Gen Zabar system, I know the finance side often doesn't stray beyond the fences of their world, and the rest of us don't stray into theirs — there's a disconnect. So I love that you're connecting finances to actual awards to determine a cost. You might also consider looking at awards and student costs in the financial aid survey.
14:44Right — and part of why I structured it this way is that IPEDS surveys all contain the institutional ID, so when you join a survey part across union'd years to a different survey, you just join on that ID and the collection year, which retains the time-series information across two surveys. IPEDS makes that clean to do with this structure. Going back to purpose — our team included a financial person and someone from the academic programs team, among others, and they're the ones who gave us these metrics and the context behind them. There are tons of IPEDS variables out there, and having people who could point at exactly which one narrows down to the metric you actually want was essential.
15:44You're hitting on something people don't always understand — we're immersed in the data world on our campus, but the truth is we don't know everything. Don't tell anybody that. It sounds like you assembled an advisory team from different areas who know the data, or at least the context, more intimately — because a number tells you how much something cost, but it doesn't tell you the context. That's why the team matters: each IPEDS survey has dozens of metrics, and nobody wants a dashboard with five hundred of them. You want the dozen or so that matter to the people using the data. But like I said, listeners, let's keep that myth going that we know everything about data — until we post this recording, anyway. So now that we have metrics, what methods did you use to select and group your peer institutions? I saw multiple groupings.
17:17That started with leadership again — whoever your stakeholders are, whoever has buy-in and will get value from the project, they had a peer set already in mind, which was a great place to start. From there we narrowed it down using Carnegie classification — we wanted other two-year, technical or community schools of similar size. Using institution size, associate's-degree-granting status, and public control narrowed us to about 400 schools — still a solid set, but clearly peers. That also let us drop some of the schools leadership had originally suggested that didn't fit. Then I used clustering: giving it attributes like fall enrollment and first-time retention — metrics that don't swing wildly year to year — places schools into a mathematical vector space and groups the ones that are close together. You don't want too many metrics; I used about five. PCA reduces those dimensions down so you can visualize it — that's the bubble chart on the poster. From that, we could see which cluster we're actually in — meaning those schools are truly like us — which clusters are nearby, which are aspirational, and which represent more short-term goals. On top of clustering, I used cosine similarity, which places schools in a similar mathematical space and measures how close one school's vector is to another's. This is something I already knew from school, and it's not the only solution — that's the nice thing about working with data, there's usually more than one way to get there.
20:45I've also created multiple peer groups, because we're usually sharing data with multiple audiences — you're at a technical school, I was at a community college, and we both have governing bodies, but our trustees always wanted to know, well, what does this look like just in Kansas? There's no single right way to build a peer group, but you do want to be intentional, and it all comes back to audience and purpose — what do you want to gain from this peer group? I love that you created an aspirational group to look toward, because that's really what this is all about: creating change, making things better for students. So let's talk about your reports — you used color in a significant way, and I'd like you to talk about that, plus anything else that made it easy for colleagues to find meaning, because there's nothing worse than handing someone a dashboard they don't understand, and then nothing happens.
22:13The color choice is fairly simple — our project lead has a marketing background and a good sense of what pops for people, so the colors were intentional: black-and-white backgrounds, with our new school color, yellow, reserved for anything we wanted to highlight — usually just WC Tech itself on a line chart, plus some button selections. Everything else stays neutral, because no matter which group or metric you're looking at, WC Tech is always in it, so it makes sense to reserve that color for that purpose. Finding meaning was probably the part that took the longest to design. At this point in the project we have time, several metrics, and several schools — three dimensions — but most visuals only handle two well. So we split it into two parts: one view looks at a single year at a time, which removes the time dimension but lets you look at multiple metrics across multiple schools, using Power BI's matrix visual for that detailed view. The rest of the pages are dedicated to yearly trends, where we removed the metric dimension instead. Rather than cramming all three dimensions into one visual, simplifying it to show just two at a time made the whole thing far more digestible for our users.
24:40For most of our colleagues who don't work in the data world every day — or who are only familiar with their little corner of it, enrollment or finance — simpler is better. I don't know if you've ever looked at what developers post on Tableau Public every day, but none of that ever really fits into my world, because the primary focus of a data strategy has to be making it approachable enough that your colleagues can create meaning from it. So again, listeners, we're going to share out all of Jon's materials — he made all of this sound easy, all the talk of matrices and cosine similarity, but you'll get the full details, and you can reach out to me or to Jon on LinkedIn if you have questions. He's the kind of person who's always willing to talk and help — that's the great thing about the data community, there aren't too many secrets, we share our love of data so everyone can grow.
26:05Thanks again for making time to talk about such an important activity — benchmarking is expected of us. All of our regional accrediting agencies expect to see comparison data, and there's really no better way than to find a data set like Jon did, one that's consistent and meaningful, and build benchmarking metrics and peer groups around it. Thank you for joining us again for another episode of Data Stakes. Data Stakes is sponsored by the Data Analytics Alliance for Higher Education — visit our website to learn more about our upcoming meetings. If you have any questions about today's conversation, don't hesitate to reach out to me at dphelps@datatelligent.ai.

Accreditation is more than just a regulatory hurdle; it’s a critical stamp of approval that validates an institution’s credibility, academic integrity, and funding eligibility. But