In this episode of the Verisian Community Podcast, we welcome Cornelia from Bayer, to discuss good programming practice in clinical programming. Cornelia highlights the critical aspects of writing good SAS code and adapting to languages such as R. Cornelia delves into how Bayer is embracing AI in clinical programming and emphasizes the pivotal role of maintaining high standards in code quality and thorough documentation, especially for regulatory compliance.
Guest: Cornelia Fulgenzi, statistical programmer at Bayer, background in biology and computer science with a specialization in SQL/databases; roughly 15 years at Bayer.
Host: Tomás Sabat Stöfsel, CEO & Co-Founder, Verisian
Topics: Cornelia's unconventional path into clinical programming, what constitutes good SAS code and documentation practices, Bayer's tiered validation model, "dinosaur code" inherited from an acquired study, regulator expectations around code (including China's agency running submitted code directly), barriers to switching from SAS to R, advice for aspiring clinical programmers, feedback on Verisian's Validation Engine
The full episode can be found on YouTube, Apple Podcasts or Spotify.
Cornelia's route into clinical programming started as a lab assistant in biology before she studied computer science, where SQL and databases became her specialty. She landed her first CRO role as a data management programmer with almost no SAS background, having only sat through a two-hour SAS presentation, but her PROC SQL fluency from her studies let her become productive within days at a CRO that happened to be PROC SQL-focused. Joining Bayer about a year and a half later, she encountered a team built around DATA step conventions rather than PROC SQL, and describes her job interview as more of a friendly technical debate, comparing DATA step and PROC SQL solutions to the same problems, than a traditional interview. Within about six months at Bayer, PROC SQL and DATA step users on her team had cross-trained each other, and she now often mixes both approaches within a single program to get the best of each.
Cornelia's definition of good code centers on three things: clear documentation so others can understand it, validation matched to risk (Bayer splits programs into standard, heavily-validated macros; high-stakes calculations that get double-programmed; and lower-risk code that gets a code-and-log review), and genuine readability. Her practical test is whether a reader can follow what a piece of code is doing without wading through dozens of lines, citing PROC SQL statements chained through three or four nested data steps as a common readability failure. Bayer's standardized environment, where study initialization, output generation, and SDTM-to-ADaM mapping steps are all handled by validated standard macros, leaves the real variation to what happens in between, which is where code quality differences between programmers actually show up.
Asked about the worst codebase she's encountered, Cornelia describes inheriting a study's programs after an acquisition: code so outdated that she had to Google techniques nobody uses anymore to understand it, with hundreds of lines doing what a modern GROUP BY or PROC SQL statement could do in 10 to 20. On the regulatory side, she notes Bayer hasn't received formal feedback criticizing code readability, but flags a structural tension: Bayer's production code relies on compiled, internal macro catalogs, and some regulators, particularly the Chinese agency, want to directly run submitted code against submitted data to confirm the results reconcile, which is harder to satisfy when core logic lives inside proprietary macros rather than the submitted program itself.
Cornelia describes herself as a "pure SAS programmer" with limited R exposure, though some Bayer colleagues are exploring R for specific steps. She identifies two real pressures toward R: it's easier to hire people who already know R than SAS, since SAS isn't widely taught at universities, and SAS licensing costs are a genuine budget consideration. But she argues the harder cost is Bayer's accumulated infrastructure, its macro catalogs, study-connection tooling, and 15+ years of deep, hard-won SAS expertise across the team, which would all need to be rebuilt, not just the code itself. She expects any real shift to be led first by CROs, who have an easier time hiring open-source-native programmers, with sponsors like Bayer following only once that shift has proven itself elsewhere. She notes tools like ChatGPT are already lowering the learning curve either direction, letting a programmer translate an R snippet to SAS (or vice versa) as a fast, practical way to pick up a second language.
Cornelia Fulgenzi's advice for newcomers is a mix of basic statistics (t-tests, p-values, the fundamentals) and general programming literacy (loops, conditionals), arguing the specific starting language matters less than having that foundation, since it makes learning the next language faster regardless of direction. She's currently taking a Coursera statistics course herself to shore up gaps from her computer science background. Asked for feedback on Verisian, she highlights the Validation Engine's ability to read through an entire large, legacy codebase without "getting tired," producing an overview of what data flows in, what transformations happen, and where variables connect, particularly valuable for surfacing dead-end datasets that are created but never used downstream, a common source of unexpected or unintended behavior in submissions.
"I would say when I really need to go through 50 lines of code to, at the end, understand what the programmer is doing, then that's probably not really a good code." — Cornelia Fulgenzi
"I would describe it like dinosaur code... I really needed to Google because no one is using it anymore to understand what the stuff is doing." — Cornelia Fulgenzi
"Especially the Chinese agency, they really want make sure that the code that you submit really belongs to the data and that everything at the end came to the same result." — Cornelia Fulgenzi
"It's not getting people with a good SAS knowledge, but people that are willing to learn quite fast and to adapt quite fast to a new environment." — Cornelia Fulgenzi
"Having a tool that just read through the whole code and don't get tired after... the first thousand lines, that's really great." — Cornelia Fulgenzi, on Verisian's Validation Engine
What does Cornelia Fulgenzi consider good clinical SAS code?
Code that is well-documented, validated at a level matched to its risk (standard macros, double programming for critical calculations, or code review for lower-risk work), and genuinely readable, meaning someone shouldn't need to trace 50 lines of nested logic to understand a simple task.
Why is legacy or "dinosaur code" a problem in clinical programming?
It's often written using outdated techniques nobody actively uses anymore, making it hard to understand or maintain, and functionally bloated — Cornelia describes rewriting hundreds of lines of inherited code down to 10 or 20 using modern SQL approaches.
Why do some regulators want to run a company's submitted SAS code directly?
To confirm the results reconcile with the submitted data. Cornelia notes this is a specific challenge for Bayer because production logic often lives inside internal, compiled macro catalogs rather than the submitted program itself, especially for agencies like China's that want to execute the code as submitted.
What's the biggest barrier to switching from SAS to R at a company like Bayer?
Not programmer skill, but sunk infrastructure: validated macro catalogs, study-connection tooling, and 15+ years of institutional SAS expertise that would all need to be rebuilt. Cornelia expects CROs, who find open-source programmers easier to hire, to lead any industry shift before large sponsors follow.
What background does Cornelia recommend for someone starting out in clinical programming?
A working foundation in basic statistics (t-tests, p-values) plus general programming literacy (loops, conditionals), regardless of which specific language someone starts with, since that foundation makes picking up additional languages faster later.
The Verisian Community Podcast brings together experts in clinical trials to exchange innovative ideas and best practices central to clinical reporting, submission and review. Aligned with Verisian's mission to accelerate the evaluation and market launch of new medical treatments, each episode features expert insights, with guests ranging from statistical programmers to medical writers, to discuss the challenges and opportunities of the latest software and technology.
You can listen to us on YouTube, Apple Podcasts or Spotify.