Four Years of the Google School of Software Engineering (SWEdu)

Delivered at the S3D Distinguished Speaker Series, Carnegie Mellon University, 17 September 2025. Introduced by David Garlan.

Here are the slides.

Below is a lightly-cleaned transcript of the talk (hesitations and false starts removed, meaning preserved). Each slide appears before the portion of the talk that discusses it. An appendix with the full audience Q&A follows the talk.


Introduction — David Garlan (host)

David Garlan: So it’s a great pleasure to welcome George back. He’s no stranger to us, of course, having done a PhD here, but also having collaborated recently with Bradley and me in teaching and revamping the software architecture course, as well as several other activities. George is very interested in connecting back with CMU and welcomes interactions with people.

The meet-and-greet filled up almost immediately, so we’ll be scheduling another day here when people can come meet with George if they want to talk to him.

I think you know what George is going to talk about. I haven’t seen the slides, so I don’t know exactly, but it will help us understand a bit about what Google does in education and software engineering. For us, a particularly interesting question is: what are the gaps George is seeing from people coming in to work at Google that we should actually be addressing in our classes? Partly what to teach, but also how to teach — George’s form of delivery within Google has a lot of innovations compared to what we’d normally think of in the traditional classroom. And the analytics and data George has gathered is really impressive; we could be doing similar things. There’s a lot to be learned. With that, George, I’ll turn it back to you.

Title slide: Four Years of the Google School of Software Engineering (SWEdu). Small Google logo top left. George Fairbanks, ghf@google.com. September 2025.

Opening & decoder ring — George

George: Thank you so much, David. It’s my distinct pleasure to be back here at CMU, seeing so many familiar faces. Hello, Bill — I didn’t get a chance to say hello. I hope not to give you a boring presentation today. Feel free to interrupt me at any time for clarification questions, because the thing I’m most worried about is dropping jargon or slang that we use inside the company. Let me give you a quick decoder ring on the very first slide: “Four Years of the Google School of Software Engineering,” which we call SWEdu, which has got to sound strange. At Google, they abbreviate the job “software engineer” into SWE. The department I’m in is EngEdu, which you could probably figure out — and we jammed those together and got SWEdu.

Two other decoder-ring items. I’ll probably say “TL” instead of spelling out “tech lead” — a software engineer who is the lead engineer for a group of people. And I’ll say “CL,” which stands for “change list.” If you’re familiar with Git, it’s exactly the same as a PR, or pull request — a proposed change to the code base. So: SWEdu, CLs, and TLs.

About the speaker slide: photo of George Fairbanks with bio text describing him as leading Google's SWEdu team, author of the book Just Enough Software Architecture, writer of the Pragmatic Designer column in IEEE Software magazine, and holder of a PhD in software engineering from Carnegie Mellon University.

A note: at 12:30 today, the other person who started this with me, Titus Winters, is going to join, so he’ll be around for the Q&A part as well. This is a guy I used to see in the mirror maybe 10 years ago. I’ve been active compared to most practitioners in getting the word out about software design and software engineering — in IEEE Software magazine, a book, and presentations, including the SATURN Software Architecture Conference that Len Bass was also a big part of. When I showed up at Google, it was a bit of a shock: in some ways they were incredibly ahead in software-engineering techniques, and yet they weren’t embracing some of the ones I thought were good ideas. That’s where all of this work began.

Agenda slide listing four sections: 01 What is SWEdu? (bolded as current), 02 What's in SWEdu?, 03 Does it work?, 04 Reflection.

Today I’ll tell you what the project is and exactly what’s inside it — because you’re probably wondering, are they teaching agile? functional programming? software architecture? — present the evidence we have about whether it works, and then share some personal reflections.

A note on process: I had to get all these slides vetted, as industrial speakers do. The things on the slides are the parts approved for publication. Especially in that last section, on reflection, you’ll be hearing much more about me, my thoughts, and what I found hard. It’s not an official statement from Google.

What is SWEdu?

What is SWEdu? slide, first bullet: The Google School of Software Engineering (SWEdu) started in 2021, founded by George Fairbanks, Titus Winters, and Kevin O'Malley.

We started this back in 2021. Kevin O’Malley was the sponsor leading the department. Titus and I found that we had expertise in software design and in software testing, and we said, look, let’s just go out there. Kevin encouraged us to shoot for the moon: the title “Google School of Software Engineering” sounds a little grand, and he said, that’s what you eventually want, isn’t it? Let’s see how close we can get.

What is SWEdu? slide, adds second bullet: SWEs learn on the job, but the lessons they learn are idiosyncratic, depending on who mentors them.

What Titus and I experienced is that engineers at Google — probably like every other company — learn on the job. But what they learn is idiosyncratic: they may be exposed to some things and not others. You may have a great lead who’s a strong mentor, or you might just have a knuckle-down-and-get-the-work kind of lead who isn’t really sharing knowledge. CMU graduates a bunch of very strong undergraduates, but it’s not clear which parts of what they don’t yet know get filled in — it’s much by happenstance.

What is SWEdu? slide, adds third bullet: self-taught lessons are idiosyncratic 'gut feel' rather than named principles, so engineers cannot easily cite or share them, leading to a Tower of Babel across the company.

They do teach themselves, but there’s a critical problem. When we talk to tech leads about these techniques and give them the names, they say: “I already know that, but I never knew somebody else had already given it a name, let alone 50 years ago.” As a result, they’re unable to share that knowledge effectively. They can say “no, no, not like that, like this,” but they don’t have a name for a principle; they can’t point to a reference. They may have reinvented information hiding all over again, but they don’t have the term to let somebody else know. Across the company you get a Tower of Babel — and I think this is pervasive in the industry. Self-taught lessons can’t be communicated very effectively.

Slide titled 'Who is SWEdu?' (title 'What is SWEdu?' shown struck through above it). Team roster: George Fairbanks (SWE), Stephanie Chiang (20% PgM), Ryan McDonough (xWF, admin and video production), Mohamed Dekhil (sponsor). Emeritus: Titus Winters, Tom Manshreck, Kevin O'Malley (sponsor), Jonathan Schuster, Bram Bout (sponsor).

Who’s been involved? The team is currently me as the only engineer, plus a part-time project manager, and someone helpful in running the logistics and editing videos. Our sponsor right now is Mohamed [Dekhil]. And there are several other people, including Titus Winters, who were key contributors.

When I started out, I had a gut feel there was a skills gap — based on having come to CMU myself and trying to absorb everything I possibly could. When I look at what everybody else was doing, I think, “I’m not sure they’ve heard of these ideas, or understand the benefit they’d get.” So one of the first things we did once we had an education process was start surveying the heck out of the tech leads — we honestly buried them in surveys: before they show up, immediately after they finish, and for a year afterwards, to see what adoption looks like.

Bar chart titled 'What is SWEdu?' showing a skills gap survey of 281 Tech Lead Seminar alumni. One bar for 'New grads': know the SWEdu ideas at 17% (bottom), a large skills-gap band, and need them at 77% (top). Link: go/swedu-lessons-learned-2025.

One of the first things that popped out: the new grads that show up don’t know all the ideas in our class — and in fact the majority of the class’s ideas are relevant even to entry-level engineers. That’s pretty good; it means we’re mostly shooting at the right target. Also, there is no way any school, even CMU, is going to get everything an engineer needs in four years. You can’t prepare somebody for a 40-year career starting from an 18-year-old; the frontier keeps moving. The implication is we’re going to need to keep teaching things in industry.

Same skills-gap bar chart, now with a second bar for 'Experienced SWEs' added alongside 'New grads': experienced SWEs know 57% and need 88% (a similarly sized gap); new grads know 17% and need 77%. Shows experienced engineers also have a substantial skills gap.

How did the experienced engineers rate? They knew a whole lot more — maybe not the terminology, but they already understood information hiding, modularity, and so on. But we were delighted to see we were still on target: even the most experienced engineers could gain things from this training.

When we started, we had a substantial amount of discussion — “discussion” is the nice word — about what the pedagogy should be. Titus and Tom had had incredible successes with a certain kind of pedagogy: weekly tip-of-the-week articles and best practices. Some of you may have heard of Google’s “Testing on the Toilet” series, where every week a new piece of paper is hung in public places inside the company with a single tip. The rules are that the tip has to have consensus and be actionable — not “let me talk to you about abstract data types,” but “this is a better library for parsing command-line flags.” Just because of that vehicle, we shy away from abstract topics there.

But I argued that when it comes to software design, there aren’t that many cut-and-dried best practices. Instead you have a bunch of ideas that, once they’re inside you, allow you to grapple with design decisions better. The variety of software Google makes is truly impressive — operating systems, device drivers, phones, telephone or essentially network switches. How would you boil down best practices for design across that range? That’s why we confronted what we’re doing here.

Can we teach best practices? slide, left column only: Software testing — Yes, easy to codify best practices; testing concepts help answer 'Have I done enough?', 'Am I doing the right kind of testing?', 'Do I have the right mix of techniques?'. Software design — No, few best practices, lots of helpful concepts.

We were guided by a distinction between education and training. I got this terminology from Tim Halloran, another person here from CMU. In the military they make a strong distinction between the two. If you need to teach someone to disassemble and rebuild a helicopter engine, that is training — there’s a standard way, everyone conforms, and you know the outcome. They also have education, which would include the war colleges. (Apologies for the military references.) There are better and worse ways of winning a war, but that’s not the same character as disassembling and reassembling an engine. We ended up leaning into the education aspect rather than the training aspect of SWEdu.

Same slide, right column added: Training vs education — training transfers right-or-wrong skills, education better equips you to wrestle with hard problems; SWEdu focuses on education but wants to add more training. The MBA metaphor — MBAs learn many lessons that aren't hard but are non-obvious, e.g. 'my revenues are bigger than my expenses, why did my company fail? Cashflow.' SWEdu teaches similar perspective shifts.

Here’s my MBA metaphor. You might naively think that as long as your company takes in more money than it costs to do stuff, it will be successful. But when you get an MBA, you learn about cash-flow analysis: what matters is when you get the money, not just how much. Companies can fail because they don’t have the right money at the right time. It’s not a profound lesson, but once you internalize it, every problem you look at shifts perspective, and you’re more likely to avoid that mistake. That is the same character as the design education we’re doing.

We have both benefits and drawbacks of doing this education in industry. The big benefit is we don’t have to test anybody — ChatGPT is not throwing out our curriculum. But my point is that evaluation for us can be as simple as giving them a survey at the end, as opposed to an academic environment, where you’re much more structured around accurately gauging how well people learn.

Here’s the flip side. If you want to become a medical doctor, the curriculum can say you’ve got to take organic chemistry — that’s just the way it works. We basically don’t have that. Every one of our students can walk out anytime they want; we have to convince them to show up in the first place, and once they’re there, convince them to stay. That influences a lot of what we’re able to do.

What is SWEdu? slide: Compatible with Google's culture. Current Google culture includes: short links ('go-links' like go/software-design), ubiquitous commenting on docs and code, and peer-review 'gates' before submitting code.

Google’s culture includes a lot of what we call “go links” — you can see one at the bottom of the page. It’s basically a small namespace: inside Google you type “go/” plus a link into a browser, kind of like Bitly or any link shortener, except the namespace is company-wide. We designed our content to fit within that go-link structure. If you’re in the middle of mentoring somebody and don’t feel like writing a whole bunch about a topic, you can say, “I know George already wrote that one-page essay — type in the go link.” We encourage these students — meaning 10-year industry veterans — to mentor other engineers via go links. We want them to understand the concept, remember the go link, and the next time the topic comes up, point them to it. The idea is that this makes the company much more efficient and starts to coordinate the vocabulary inside the company.

Same slide, right column added: Mentoring via go-links — write short single-topic pages with go-links, teach SWEs the concepts, SWEs mentor each other citing the go-links, effective within the flow of work. Recent LLM changes — internal LLMs have ingested our content and are increasingly citing our content.

As everybody knows, LLMs are a big deal these days. One of the most interesting things I’ve been seeing in the last couple of months is that our internal LLMs have now been trained on the content we’ve created. If you ask them anything about software design or software testing, they’re incredibly likely to cite our content — which is delightful to see. So we’re not only training engineers, we’re training the bots.

Comic-style slide titled 'SWEdu content helps your team's workflow.' A stick figure labeled SWE asks a stick figure labeled Tech Lead, 'TL, will you review my CL / design doc?' The Tech Lead thinks 'Hmm, looks like conceptual trouble... And I'm already overbooked today.' Caption: TL recognizes SWE with a conceptual problem. Footnote: CL = change list = pull request.

Here’s a short sequence of slides that puts, in a nutshell, how we hope to get ideas into the company. Imagine you’re a tech lead reviewing some work — a design document or a proposed change, a CL. Somebody says, “I’d like to get your input on this.” We have a strong culture of doing that. If it’s a trivial change — you forgot a semicolon or you have a typo — no problem, you just point that out. But as soon as you say the equivalent of “I think your characters in your story lack motivation; we need to talk about the fundamentals of storytelling,” then you’re in trouble as a mentor. You only have a few options.

Comic slide 'Option 1: Teach on the fly in CL reviews.' SWE asks the Tech Lead to review a CL/design doc; TL says 'Sure!' and thinks 'Looks like I'm teaching concepts via CL comments, yet again.' Caption: Across Google, vast time wasted as TLs write and re-write explanations.

Option one: teach on the fly — you try to type that essay into the comments section of a document. But it’s an incredible burden on the leader, and often it doesn’t happen because of the time burden.

Comic slide 'Option 2: Individual live mentoring.' TL says 'Let's sit down to discuss the ideas behind your code'; SWE says 'Cool, I'll learn a lot from that'; TL thinks 'I'm robbing Peter to pay Paul...' Caption: Inefficient, TLs often starved for time.

Option two: walk over to somebody’s desk and say, “let’s chat through this; I was surprised to see this proposal.” Again, that’s incredibly expensive.

Comic slide 'Option 3: Code it yourself.' TL says 'No, write the code like this'; SWE replies sheepishly 'Thanks'; TL thinks 'Why can't others do it like I can?' Caption: SWE learns concepts slowly, or not at all.

Option three: say “nope, not like that, like this” and hand them the solution. That’s not the best for teaching, and culturally you feel sheepish afterwards — “my TL knows how to do this, but I’m not very good at it.”

Comic slide 'Option 4: Read a textbook.' TL says 'Go read this book' and thinks 'This rarely works, but I'm too busy'; SWE replies 'Um, this is due tomorrow.' Caption: Books are great, but hard to use in-the-moment.

Option four: maybe you learned this idea from a textbook, and even remember which one. But if you’re in the middle of a review cycle — here are 20 lines of code, can you take a look? — the last thing you want to do is insert “please read textbook here before continuing.” It’s just impractical.

Comic slide 'Option 5: On-demand modular content.' TL replies 'Try reading go/example... then let's chat'; SWE says 'Sounds good'; TL thinks 'This saves some time... Thank you SWEdu!' Caption: TLs use pre-packaged explanations to teach SWEs concepts, at their time of need.

So what we ended up with is building a whole bunch of short pages, putting them behind go links, and getting them out to the people doing the reviews — the tech leads. The idea is that this saves everybody time and starts to pull the company together.

How do we train TLs slide. Flagship course: the SWEdu Tech Lead (TL) Seminar. Audience: 281 Google/Alphabet TLs. Duration: two half-days per week for 6 weeks. Format: graduate seminar with lots of pre-reading. Model: train-the-mentor. Evaluation: surveys. We've run 8 cohorts; this is how we 'prime the pump.' Right side: screenshot of the internal go/swedu-tl-seminar website describing live instruction and self-study options.

How do we train these people? These are the tech leads. It’s a six-week program, two half-days per week. Typically I do Tuesdays on design and Titus does Thursdays on testing. It’s the format of a graduate seminar, so before each class there are a handful of readings they’ve done ahead of time. (At the end of the slide deck you’ll have available are the readings we’ve assigned.) The whole model is to influence these technical leaders and send them back into the company — effectively a train-the-trainer, or train-the-mentor, situation. We want to encourage mentoring and make them strong, effective mentors.

SWEdu's multi-year game plan slide, part 1: Create content (and link to existing content) — written website go/software-design, video courses (Design by Contract, SWEdu TL Seminar), video podcasts and tech talks. Nurture a community — mailing list and chat group, TL Seminar alumni groups.

Our multi-year game plan is to do a whole bunch of things that, summed together, will change Google’s culture and improve the state of the practice. First, we have to create content. If you think about many core ideas in software engineering, there isn’t a single essay you can point to that reveals that idea. We have to write those essays — the one-pager on that thing. It might be in the middle of Len’s book on software architecture, but they’re not going to read the first six chapters to get to chapter seven, get the idea, and get back to work.

Same slide, adds: Prime the pump (train the mentors, i.e. the TLs; focus community attention) and Patiently wait for geometric growth (ideas flow from TLs to teams to all SWEs), illustrated with a four-panel 'Gru's Plan' meme captioned 'Ideas to TLs,' 'TLs to teams,' 'Teams to all SWEs,' and 'Patiently wait?' (the presenter's head-scratch panel).

Second, we try to nurture a community. If we’re not constantly drawing attention to these topics, people shift their attention elsewhere. There’s a renewing of interest. We’re “priming the pump” by teaching all these folks, then patiently waiting for change in the company. If you’re familiar with this meme template, the “waiting patiently” part is the hard part. We’re starting to see real signs of change, but we wish it had happened already and that we were farther along.

Screenshot of the internal go/software-design wiki page 'Quality attribute priority,' dated 2025-01-27. Summary box: 'A quality attribute priority expresses the relative priority of several quality attributes.' Example: 'Scalability > Usability > Latency > Modifiability.' Section 'Tradeoffs are inevitable' explains that systems can't have every desirable quality and thinking about priorities helps with tradeoffs.

What does one of these go-linked pages look like? Here’s one on quality-attribute priorities — probably a bit of an eye chart, but you’ll have the slides later. In the blue box is a summary: “A quality attribute priority expresses the relative priority of several quality attributes.” It’s an idea you might need to reference, so you can go to go/quality-attribute-priority and drop that link when you’re mentoring anyone.

Agenda slide with '02 What's in SWEdu?' now bolded as current section.

What's in SWEdu? slide: SWEdu consists of written materials (a website), courses (Design by Contract and the flagship Tech Lead Seminar), recorded and live videos, and community (video podcast, chat, email list). Link: go/swedu-tl-seminar.

So what’s the content, exactly? Our project consists of four things: written materials, courses, recorded and live videos, and a community. On written materials: we have a guide on software architecture (you saw one of the pages), plus a collection of miscellaneous design concepts that haven’t made their way into one specific guide. We have a Design by Contract guide, because we find DBC ideas are very compatible with testing — code takes on a contractual nature. If you’re a functional-programming fan, you realize this is like the gateway drug for procedural programmers. We also have a guide on design diagrams and diagramming tools, which as it turns out is the most popular thing we’ve done. Plus a guide on error handling, and a whole bunch of tech talks and video podcasts on the website.

Same slide with 'Written materials' highlighted; right box lists design-concept vocabulary covered: quality as a strategy, functionality vs quality attributes/tradeoffs, design for testability/test-size tradeoffs, OODA loop (quick feedback, shift-left), design by contract, intellectual and statistical control, stable code (E-type and S-type, stable sub-problems), flaky vs brittle tests, and actionable test failures.

Same slide with 'Courses' highlighted; right box: the flagship Tech Lead Seminar is a 6-week course whose content is organized into Small, Medium, and Large buckets plus cross-cutting themes.

Here’s what’s in the courses: a Design by Contract course, and the flagship course I keep referring to — the Tech Lead Seminar, the six-week one. We sort of force-fit the content into small, medium, and large because we need a structure for the class.

Same slide, right box now shows 'Small' topic details: quality as a strategy ('high quality bricks'), functionality vs quality attributes/tradeoffs, design for testability/test size tradeoffs, OODA loop (quick feedback, shift-left), design by contract, intellectual and statistical control, stable code (E-type and S-type, stable sub-problems), flaky vs brittle tests/actionable test failures.

In “small”: the very first topic is quality as a strategy. The idea is that Google started out making a high-quality distributed system from low-quality — i.e., unreliable — PCs. That was a wild innovation. But think of it this way: it’s a lot easier to build a distributed system out of high-quality parts than low-quality parts. If you can make stuff good, you have an easier time building a high wall. That’s the metaphor we use the whole time.

Same slide, right box shows 'Medium' topic details: modules and coupling / complexity reduction, test doubles (fakes, stubs, mocks), error handling and typeful programming, fuzzing and property-based testing.

In “medium”: modules and coupling, error-handling topics, and various kinds of test doubles.

Same slide, right box shows 'Large' topic details: software development processes (waterfall, incremental, iterative), ur-technical debt, integration tests, fidelity/speed/cost tradeoffs, software architecture, continuous integration.

In “large”: a very brief overview of software development processes, where tech debt comes from, and continuous integration.

Same slide, right box shows 'Cross-cutting themes': intellectual vs statistical control, quality enables velocity / 'quality is free', virtuous and vicious cycles, SWE growth and mentoring, multiple perspectives, shift-left/pulled-right, seeking balance, signal processing, OODA loop (observe orient decide act), accidental and essential complexity.

After teaching the course a couple of times, we realized that during class discussions — because it’s run as a seminar — several cross-cutting themes kept coming up that were never one specific topic. One is this idea of intellectual and statistical control. As everyone knows, industry has gotten onto the testing bandwagon — that’s an example of statistical control: you look for specific cases and test them before you ship. You can imagine a factory spot-checking its products. We encourage people to do the thing every academic here is intimately familiar with: you should be able to think through your software, not just have empirical evidence that it works. We want you to have both. But it’s almost a radical idea at this point — people have so embraced empiricism that they’ve kind of abandoned the idea they might be able to think through whether their software works. Getting that on the table is a big part of these cross-cutting themes.

Does it work?

Agenda slide with '03 Does it work?' now bolded as current section.

Does it work? slide: 'Yes, but we'd like to do better.' Let's dig into three areas: Effectiveness, Pedagogy, Adoption (no highlight yet).

OK, here’s where the fun stuff begins: does it work? The answer is yes — it’s working, but we’d like to do better. Let me dig into three topics: effectiveness, pedagogy, and adoption.

Does it work? slide with 'Effectiveness' highlighted. Right box: '2021 Tech Lead Seminar Applicants (pre-training)' quotes: 'SWEdu design/testing ideas will save each member of my team an average of X weeks per year of effort' — Average 5.6 wks/yr/SWE, StDev not shown; 'SWEdu design/testing ideas will accelerate my team by X%' — Average shown as a large positive percentage. Link: go/swedu-lessons-learned-2025.

On effectiveness: before we did any of this training, we ran a survey asking about hypothetical training on these topics — how much time would it save your entry-level programmers? L3s and L4s are the entry- and mid-level ladders. What we got back was it would save them five to six weeks per year — a pretty incredible number. What was even more impressive: the more seniority, the more years of experience, the higher your level, or whether you’re in a leadership or management position, the higher the number. That’s great — it indicates people besides me were perceiving there was efficiency to be gained.

Does it work? slide with 'Effectiveness' highlighted. Right box: '2021-2025 TL Seminar Alumni (post-training)' quotes: 'SWEdu design / testing ideas will save each member of my team an average of X weeks per year of effort' — Average 8.0 wks/yr/SWE, StDev 7.2 weeks; 'SWEdu design / testing ideas will accelerate my team by X%' — Average 25.7%, StDev 18%.

We did a similar survey after delivering the training, which has the benefit of being after the fact — they’ve actually seen the content. Now we can ask, how much will this content help your engineers? (With the caveat that they’re predicting productivity — I wish Ciera Jaspan were here to talk about how hard it is to measure and predict productivity.) What we found was a pretty incredible number: every member of their team, as a result of the tech lead taking the training, would be about two months per year more efficient — with a standard deviation that’s off the charts. That kind of made sense to me, because we have a buffet of ideas and say, “any of these any good?” Some people say, “yeah, our team needs this testing idea”; some say “I wish I knew this design idea last year, because we just made a bunch of mistakes as a result.” So some people end up with very large numbers because they now recognize a mistake they’d had. But the number has been very consistent, despite the large distribution. And if you believe in the wisdom of crowds, we got data from 280 tech leads.

Same post-training quote box as before (8.0 wks/yr/SWE average, 25.7% acceleration average), now with an added red '15x ROI' callout bubble pointing at the numbers.

All in all, we got three different predictions: the one I showed before training; this one about time saved; and a third where we asked them to estimate the acceleration on their team. All are quite large in any case. One of our partners — we train a lot of Waymo engineers — did a back-of-the-envelope calculation: you’re investing on the order of 12 hours per week for your tech lead over six weeks, and you’re telling me you’ll get eight weeks per year out of your engineers? That’s like a 15x ROI in terms of time invested and time saved.

Does it work? slide with 'Effectiveness' highlighted. Right box titled 'Boosts SWE morale & opinion of Google': quote 'Attending SWEdu made me happier' — Likert average 4.3/5, StDev 0.9; quote 'Attending SWEdu improved my opinion of Google as an employer' — Likert average 4.3/5, StDev 0.8. Below, two horizontal stacked-bar Likert charts (Strongly disagree to Strongly agree) for the same two statements, both showing responses concentrated heavily in Agree/Strongly agree.

We were delighted to find that people loved this class. It actually made them happier and improved their opinion of their employer. So if you let employees take this training, they come out happier and more productive.

Does it work? slide with 'Pedagogy' highlighted. Right box (paraphrased survey quotes): appreciation for identifying concepts and providing written references to cite, and that the common vocabulary dramatically reduces the Tower of Babel effect.

Shifting to pedagogy: for this to be effective, we need both the training materials — the class — and the written materials. Because the whole idea is to bring tech leads in, talk about topics, and make them effective mentors. The effective-mentor part requires writing all those web pages. We hear this in the surveys: “It was great that you identified the concepts, and I really need those written references so I can cite them.” They also overwhelmingly point out that the common vocabulary dramatically reduces the Tower of Babel effect — engineers communicate with precise terminology.

Does it work? slide with 'Pedagogy' highlighted. Right box describes a controlled experiment: 90 candidates split into three groups of 30 — a control group, a written-materials-only group, and a live-class-plus-materials group. The written-materials-only group disengaged ('ghosted') within about two weeks.

One thing we were dismayed by is an experiment we set up the very first time we ran this. We carefully got 90 candidates and divided them into three groups of 30: a control group; a group that got just the written materials but was not allowed to participate live; and a group that got the whole shebang — live class and all materials. We set it up to argue to management that it’s really effective to send people to live training. What we found: all groups were very excited at the beginning — at least the live group and the written-materials group. But after two weeks, every single person in the written-materials group ghosted us. They weren’t operating at a lower level; they just stopped the training entirely. Two weeks ago, very excited; two weeks later, we can’t get them to answer an email. We were sort of shocked — though it comes across as an obvious conclusion in hindsight.

You’ve got extremely busy people who always have something they need to be doing today. If you don’t have something on the calendar and part of a moving train — you’ve got to keep up, do the readings, show up Tuesdays and Thursdays — it’s like a gym buddy. There’s a reason these patterns exist for human beings. So our conclusion: do not consider making these videos and the training available on our website to be a solution. You need to continue to schedule classes and push this thing; otherwise there’s always something else you should be doing.

Does it work? slide with 'Adoption' highlighted. Right box: 'Teams are adopting SWEdu ideas slowly. Most are unaware SWEdu exists.' Quotes: 'TLs need support to enact change' — 'Keep spreading the knowledge about fundamentals. It's hard for me alone to teach this to my team. The reinforcement from broader context helps build momentum.' 'Need more & better mentoring materials' — 'It's important to have written materials and documentation to increase the likelihood of getting team buy-in.'

On the last topic: are teams adopting the ideas? Yes, they are — but we’re still running into a marketing and publicity problem. Most engineers in the company aren’t aware we exist. In the last group, somebody said, “George, I know about your work because I read it when I was at ThoughtWorks, and I didn’t know you were at this company until I got the announcement for this course.” It’s a big company; just getting attention is very hard.

They also tell us it’s not just having the materials — having them on a website gives the TL credibility when they argue for a position. “Hey, I think we should do this — oh, by the way, the Google standard is over here” — that’s way more persuasive than “hey, that’s your idea.” We saw a pretty good adoption rate across the different cohorts. Six cohorts have completed all the surveying; there are two more cohorts within the past year that haven’t finished all the surveying.

What I haven’t mentioned is that the first four cohorts, one through four, were taught live by myself and Titus. Then starting with five, six, and seven, they replayed the videos, because Titus was at a different company by then. That actually works OK — not as well; people aren’t as delighted, but they still come out with pretty positive stuff.

Does it work? slide with 'Adoption' highlighted. A bar chart titled 'Adoption Rate by Cohort' (an 'eye chart' of many software-design topics as the x-axis, colored stacked bars per cohort showing percentage of responses), illustrating that some design topics were highly adopted while others were barely adopted.

Does it work? slide with 'Adoption' highlighted. Stacked bar chart 'Adoption Status by Design Topic': for each design topic on the x-axis (terminology, tradeoffs, etc.), a 100%-stacked bar broken into five response categories (my team and I think this way now after SWEdu; my team and I thought this way before SWEdu; I now think this way but my team does not; I thought this way before but my team did not; my team and I do not think this way now), color-coded dark green/light green/purple/light blue/pink, showing wide variation in adoption across topics.

What we found — this is an eye chart, look at the pretty colors — these are the various topics for software design. Some topics were highly adopted; some were barely adopted. That’s what I want you to take away. And here are the testing topics: not quite as much of a drop-off — a bit more consistency. That’s probably because the idea of doing software testing at Google has a 15-year head start compared to the design ideas at Google.

Does it work? slide with 'Adoption' highlighted. Stacked bar chart 'Adoption Status by Testing Topic' for six testing topics (terminology flaky vs. brittle, properties of good tests, connecting test cases to DBC, testing as signal-processing/CI-as-alerting, property-based testing, fuzzing), same five-category color coding as the design-topic chart. Caption: Testing topics were more consistent.

Charlie Garrod (audience): Can you please clarify how you’re measuring whatever metric you’re actually talking about?

George: Yeah, it’s actually a bit complicated. When you look at the slides afterwards, in the top-right corner, the decoder ring: the green says “my team and I think about testing this way”; the lighter green says “my team and I thought about testing this way before”; purple is “I now think about testing this way”; blue is “I thought about testing this way before, but my team did not.” There’s an equivalent version of this, and we track it immediately after class and then quarterly for a year, and we see the numbers creep up a bit.

If you think about it from those last two slides — are you successful, George, in getting your curriculum into the company? — the answer is no. If you think about it a different way — we have a buffet of ideas; we can’t possibly come up with one curriculum that works for device drivers and backends and frontends and phones and you name it — then I think we’re doing OK, because everybody seems to find something to eat at the buffet.

Does it work? slide with 'Adoption' highlighted. Table 'Number of topics adopted' vs '% of TLs adopting': 5+ topics 99%, 6+ 98%, 7+ 96%, 8+ 91%, 9+ 87%, 10+ 82%, 11+ 77%, 12+ 68%, 13+ 58%, 14+ 46%, 15+ 34%, 16+ 27%, 17+ 21%, all 18 surveyed topics 8%. Caption: Overall, everyone found some ideas to adopt.

Essentially everybody comes away with multiple topics they’re excited about. One thing we do is deliberately teach some topics adjacent to code — immediately actionable by the team — and we see the best adoption rates for those.

Does it work? slide with 'Adoption' highlighted. Table of Subject/SWEdu Topic/% of TLs adopting, Design and Testing rows. Design topics highlighted: Error Handling 94%, Prefactoring 91%. Other Design rows: Information hiding (modularity) 90%, Intellectual Control 89%, Design Process 87%, Quality Attributes 85%, Design by Contract 83%, Architecture Styles 58%, ADRs 52%, Views 48%, Connectors 42%, Architecture decision template 40%. Testing rows: Properties of good tests 98%, Terminology flaky vs brittle 96%, Connecting test cases to DBC 82%, Testing as signal-processing/CI-as-alerting 81%, Property-based testing 78%, Fuzzing 68%.

We see the worst adoption rates for the things I’m most passionate about, and that’s an embarrassment. (David and Mary are going to get after me later.) But in some ways it’s understandable: those ideas have the least penetration into the industry, they’re the least mainstream at this point, and they’re the most abstract. I have four hours to cover these topics, so maybe it’s not surprising they’re not fully adopted.

Same table, now with the lowest-adopted Design rows highlighted instead: Architecture Styles 58%, ADRs 52%, Views 48%, Connectors 42%, Architecture decision template 40%. Caption: 'We only have 4 hours to cover software architecture and it's a big & abstract topic, so perhaps it's not surprising that the lowest adoption is on those topics.'

As for viral adoption, based on a very small survey and some individual interviews: it seems about 70% of the ideas make their way into tech leads, about a quarter of that jumps over to the team, and we’re not sure how many jump into the rest of Google. I’d love those numbers to be higher, but at least things are moving — which is awfully good.

Does it work? slide with 'Adoption' highlighted. Right box: 'SWEdu isn't yet viral. The ideas are solid, they transfer to TLs, but not yet Google-wide.' Our strategy: grassroots / viral spread. SWEdu to TLs: 70% transfer rate. TLs to teams: 24% transfer rate. ...to all SWEs: ?? transfer rate. Caption: Numbers based on a small survey.

Leading up to this presentation, I realized Google Analytics is hooked up to our internal websites too, so I pulled the numbers: how many people looked at go/software-design in the past year? I was honestly shocked: 74,000 people read what I wrote last year. I’m still letting that wash over me — here I am, some guy, saying you should know about quality attributes and tradeoffs, and now there are go links for them. I feel very encouraged this is actually working when I see a number like that. I assume it’s not product managers reading our website; I assume it’s engineers — but we don’t actually have that data.

Does it work? slide with 'Adoption' highlighted. Screenshot of Google Analytics dashboard for the SWEdu property: Active users 74K (up 558.1%), Event count 628K (up 579.4%), Key events 0, Views 188K (up 592.9%), over the last 12 months, with a line chart of active users climbing from about 1K to a peak near 3K. Caption: Most SWEs used our g3docs in the past year.

So here’s a summary I’d like to leave as a capstone. One alum says: these ideas are good; the importance of saturation cannot be underestimated. It is not one-and-done — it’s not that George or Titus said the right thing in class and now the company has changed. It never works like that. And going back to you: it’s impossible to imagine you could say the right thing to an undergraduate and solve industrial software engineering. It requires everyone to keep repeating the good ideas, saturating them. The more people you hear a good idea from — “hey, I really think we should test this,” “hey, I really think it’d be great if we had an architecture model” — the more practices change. If it’s said once, it drops on the floor and nobody’s behavior changes.

Does it work? slide with 'Adoption' highlighted. Right box, quote from a SWEdu alum: 'I want to underline that point about saturation [of SWEdu ideas as important for adoption] ... even after doing some short series of talks on some of the concepts to the team, as [the team evolves] there's continual education. The other team members just don't necessarily absorb everything as fully as you might expect from a one hour talk, so it's repeated education. ... It's a lot of educational demand relative to the need and [the need for] higher saturation.'

Does it work? slide showing a horizontal stacked-bar Likert chart 'Number of Respondents' (scale -250 to 250) for five statements: 'SWEdu TL Seminar was worth my time,' 'I learned something valuable from another SWEdu TL Seminar TL,' 'Attending SWEdu TL Seminar made me happier,' 'Attending SWEdu TL Seminar improved my opinion of Google/Alphabet as an employer,' and 'I would recommend SWEdu TL Seminar to other TLs.' All five bars are dominated by Agree and Strongly Agree (blue) responses, with small Neutral/Disagree segments.

As a capstone for “does it work”: these are Likert-scale results — was it worth my time, I learned something valuable from another tech lead, it made me happier, attending improved my opinion of the company, I’d recommend it to others. All I really want you to see is that it’s all over on one side; in general, people leave this course very happy. And that’s after bombarding them with surveys — and every single one of them doesn’t carve out enough time. They always think, “I’m a very efficient person; I can sneak this course into my regular schedule” — and they complain nonstop. I say, “go back; we warn you every single day — you need to carve time out; we make you get permission from your manager.” I think it takes that much time.

Does it work? slide continuing the Adoption theme, showing a stat box comparing SWEdu to other trainings TLs have taken, headlined with the finding that the median percentile rating was the 90th percentile, and about a quarter of respondents rated it the 100th percentile. Link: go/swedu-lessons-learned-2025.

When people come out of this, the median percentile they give for how it rates compared to all the other education they’ve done is the 90th percentile — which ain’t bad. But what’s even more amazing is that a quarter of the people said literally the 100th percentile, which I interpret to mean a quarter of the people think this is the best thing they’ve ever had in their life.

Personal reflection

Summary slide recapping the 'Does it work?' section as a numbered list of roughly ten lessons learned spanning effectiveness, pedagogy, and adoption — including that SWEdu boosts productivity and morale, that written materials alone don't sustain engagement, that SWEdu isn't yet viral company-wide, and that saturation/repetition is necessary for adoption.

Agenda slide with '04 Reflection' now bolded as current section.

So now I shift to my personal reflection, and I’ll try to connect it to some things relevant to the handoff and the relationship between academic and industrial education. First: I’ve been to other presentations where someone at an engineering company says “I ran this project, we trained these people, we had some good results,” and I’m in awe that they got anything to work — because having tried to do it myself, I realize so many things have to come together. I’m incredibly grateful to the people who sponsored this project, because it’s hard — people ask, “why are you spending scarce resources on this, why don’t you write more code?” It also relies on a great number of good ideas — which I’ve liberally stolen from everyone here at Carnegie Mellon — and on soft skills: finding a way not to irritate people and to be persuasive to get where you want.

Reflection slide, first lesson: grassroots efforts feel like a tiny boat in the ocean — SWEdu is one small team trying to shift a company of tens of thousands of engineers, relying on scarce sponsorship and goodwill rather than top-down mandate.

Another thing: many of the lessons we’re teaching are disruptive. There’s plenty of education that’s incremental, and I think everyone wants that model — “hey, here’s a great way to parse command-line arguments; I don’t have to throw away any ideas I hold dear; there’s a new library, it saved me 10 seconds, great.” But many of the ideas we’re talking about say: hold on, you’re doing pretty well, but you need to back up, get rid of some bad habits in order to do even stronger. And what’s more, you now need to convince other people to do the same. That is hard. If you tell a team they’re not testing the right way, or they’ve been undervaluing software design, that’s a disruptive change.

Reflection slide, second lesson added: additive changes (new libraries, new tips) are easy to adopt, but many SWEdu ideas are disruptive — they require unlearning habits and convincing teammates, not just adding a new trick.

Titus Winters used to say that we stumbled upon teaching design and testing in a way that was very convenient — they fit together like learning calculus and physics. One is the more abstract version; one is the more applied version. You can get an intuition for a phenomenon over here, and understand it in general over here. Being able to alternate those two was incredibly valuable in making these lessons stick. I’ve become a big fan of making sure you have the concrete along with the abstract.

Reflection slide, third lesson added: design and testing complement each other like calculus and physics — one abstract, one applied — and alternating between the two makes lessons stick.

We have a difficult problem: we aren’t doing skills transfer — we’re doing education. We need to send these tech leads back to their teams. Assume a team is at a local maximum of practices; they’re doing a good job. Now we’ve convinced the TL the team could be up here somewhere, but to get there they need to start doing things in a worse way until they can put everything back together in a better way. So we really need to train these mentors and teachers to confront arguments and objections from their team that we can’t possibly predict. It’s a very difficult thing to prepare someone to do.

Reflection slide, fourth lesson added: teams are at a local maximum of practice; moving to a better local maximum requires temporarily getting worse, and TLs need to be equipped to handle their team's objections along the way — a 'skills transfer' problem SWEdu doesn't yet fully solve.

Finally, I think there may be an expectation in academia about what the baseline practices in industry are. I’m not just talking about Google — I’ve been a consultant at a bunch of different companies. The median code basically looks like it’s from 1970. That’s not necessarily bad — you can make money writing 1970 code. But if you think you’ve taught them abstract data types so they’ll use them, they’ll use them in the standard library — they won’t necessarily write their own. They may never go, “wait a second, the big idea behind abstract data types is much bigger than lists, sets, and queues.” So practice changes very slowly. There’s a lot of work to do to actually improve the state of the practice, which is wildly uneven — some teams are way ahead, and some are still writing mundane stuff you could write in BASIC.

Reflection slide, fifth lesson added: the baseline practice in industry is lower than academia assumes — much production code resembles 1970s style — and adoption of a taught concept (like abstract data types) is often shallow, limited to using the standard library rather than internalizing the underlying idea.

So how can we improve? First, I think this kind of training needs to include both skills training and education. We do an OK job on the conceptual education part; we do not do a great job on “here’s how you disassemble and reassemble the helicopter.” When a tech lead goes back to their team, we’re essentially asking them to create educational materials or mentor without anything except the slides we gave them — and those slides were geared at an experienced audience. When they’ve got early-career people, it’s inappropriate to ask them to do all that mentoring without skills training.

Reflection slide: How SWEdu can improve. Add skills training — alumni TLs struggle to mentor their team; some ideas can be training (not education), e.g. Design by Contract, modularity, error handling. Leadership training — leaders need the terminology and concepts too, they balance short- and long-term goals, and they control the purse.

Second, a lot of tech leads come back and say, “not only do I need to convince my team, I need to convince my management to see this the way I now see it.” Third, we need better channels. As I mentioned, not everybody inside the company knows we exist. Strangely enough, the best way we advertise what’s going on is an email newsletter that goes out once a month on education efforts, and we get thousands of people looking at our website as a result. Google is way too big right now for the guerrilla tactics the Testing on the Toilet folks used — they had a wild idea and started putting posters up. You can’t do that within a company like we are now.

Same Reflection slide, adds a second column. Better channels — best channel today is the EngEdu newsletter; Google is too big for Testing-on-the-Toilet-style guerrilla marketing. Wider expertise — George Fairbanks (ghf@) plus Titus Winters (titus@) had fewer knowledge gaps together than George alone has today; it's hard to find candidates with breadth like Jonathan Schuster (jschust@).

The last point: we’re suffering from limited expertise. I’ve done everything I can to soak up good ideas like a sponge, but I think everyone here realizes there are so many other things you just don’t know. I could never have created this course without Titus — he brought an incredible amount of knowledge, chose the topics, and explained them extremely well. We still rely on his lecturing for that. Jonathan Schuster, who has a PhD from Northwestern, brought in a bunch of programming-language expertise. I wish we had more of that. I’m not trying to say this is the George project — this only works because we have these other folks who filled in the gaps in my skill gap.

Reflection slide: 'TLs say we're on target.' Gemini summary of sentiment from alumni written feedback (two columns of quotes): participants overwhelmingly commend the seminar for providing a 'lingua franca' formalizing previously vague concepts; a pervasive theme is the universal relevance of SWEdu content; 'intellectual control' is consistently highlighted as a valuable mental model; Design by Contract and error handling are found among the most helpful, directly applicable topics.

Overall, the TLs say we’re on target. The neat thing is now that we have these AI engines, I can take every bit of written content, put it into a text file, feed it to Gemini, and say “please don’t butter me up — be as neutral as possible and tell me what they said about these topics.” So here are quotes from Gemini’s summary of the alumni written feedback: participants overwhelmingly appreciated the common language; they thought the ideas were universally relevant; they thought intellectual control was something they’d been missing; and they thought the ideas were applicable to their day-to-day work. They love the content and think it’s useful in a world where we’re using LLMs more and more.

Reflection slide: 'TLs say LLMs + SWEdu = ❤.' Gemini summary quote: 'There is a strong consensus that SWEdu's principles, particularly Design by Contract, are directly applicable and even more important in this new [LLM] paradigm for ensuring quality and maintaining intellectual control over AI-generated code.' Below, a horizontal stacked-bar Likert chart (scale -25 to 25) for two statements, 'SWEs need SWEdu design ideas in a future with LLMs' and 'SWEs need SWEdu testing ideas in a future with LLMs,' both dominated by Strongly Agree responses.

This is a subject of active debate — you’ve heard lots of opinions about what the future holds. Are we heading to a future where requirements people merely dictate their requirements, engage in some Q&A, and the code pops out? Nobody knows. But what the TLs tell us is that as a result of this training, they understand the world of software development better, and they feel that’s a necessary skill to guide the LLMs and use them successfully. With that, I want to thank you all again. I couldn’t have gotten here without many of the people in this room, and I certainly couldn’t have gotten there without the people who helped create this course. Titus Winters, I think, may be listening in — thank you, Titus. With that, let’s have questions.

Closing slide, a repeat of the title slide (Four Years of the Google School of Software Engineering (SWEdu), George Fairbanks, September 2025), left on screen through the Q&A.


Discussion

Audience: Thank you for the talk. In one of your slides, you mention diagrams and diagramming. With my students, I’ve found they have a problem expressing themselves pictorially — whether it’s UML, arrows, whatever. What is the problem you refer to? What’s on those pages?

George: I can tell you what’s on those pages. First is a table listing a bunch of different diagramming tools with a summary of how they can be used. For example: can I edit the source code of the diagram, or is it always me visually moving things around? Can I drop it into Markdown — we have work documents based on Markdown? Can I drop it into Google Docs? That seems very mundane. The other pages are my attempt to get people to recognize that a diagram is a model. The reason you’re having trouble drawing the diagram is that the model is unclear to you — not because you don’t understand boxes and lines. Boxes and lines are simple; models are surprisingly complicated. I don’t know how well that lesson lands, but I have a feeling engineers, because they’re guided through the design-doc template (“put diagram here”), go “OK, tell me about diagrams.” So it ends up being a very popular page.

Michael (audience): I’ve experienced it. We used the Google Design Doc template based on Titus’s book in our class, and we had exactly the same phenomenon. Although with LLMs, the diagrams have gotten a lot prettier but a lot more raw — that’s just been my experience. But I had an actual question. You mentioned at the beginning this “17% vs. 77%” skills-gap figure in intro software engineering, and you also talk about the value of providing a lingua franca — people recognize a concept but just didn’t have the terminology. You also had to give us a bunch of terminology that we as experts didn’t know. So my question is: A) how much of this is teaching “the Google way,” and how much is arriving at an agreed-upon vocabulary? How much is vocabulary and how much is conceptual learning? Having a shared vocabulary is probably good — I’m not criticizing — but what’s your sense of the balance?

George: I think it’s a very fair question. Titus, are you on the call? Do you want to take this one?

Titus Winters (remote): Well, actually, because you’re outside the company, you can say whatever you want, and you’re probably in a better position to chat about the vocabulary or the concepts that I’m bringing in for software design.

Yeah, I think it’s really important to remember: you can’t give someone a solution to a problem they don’t realize they have. Also, it’s a lot easier to put a name on an experience they’ve already had than it is to describe in the abstract a scenario they may experience in the future and give them a name for it — and then they take a test two weeks later, don’t think about it again until they encounter the problem. How would we expect them to name that correctly five or ten years later? One of the huge insights for me from this experience is that vocabulary is an incredibly powerful thing to focus on, specifically for early-to-mid-career engineers — like five-year-experience engineers — because they have so many working hours under their belt and a breadth of experience that’s just not available in the classroom. At that point, vocabulary becomes so much more relevant.

I’ve also spoken in the last year about a cultural phenomenon I see in tech — I can’t swear to it, but it all kind of lines up in my head: because of psychological-safety concerns, fitting-in concerns, and maybe just being polite and not interrupting the flow of conversation, our industry doesn’t stop to ask “hey, what does that word mean?” You just pick it up through context, and the person who just said it picked it up through context too. This is why words don’t actually mean anything in industrial spaces. The number of times I’ve had to ask “what does continuous integration mean?” terrifies me. Same applies in testing, design, and all these things, because we have an incredibly squishy understanding — no one has been given vocabulary at a time in their career when that vocabulary would stick. Does that answer your question, Michael, I assume?

Michael (audience): I mean, what I’m hearing is that the vocabulary is really important, but it’s not that they have a non-Google-specific word for this concept and you’re centralizing on the Google vocabulary. It’s that they’ve encountered the concept but don’t have a word to put to it — which is a different way of thinking about it.

Titus Winters (remote): I think it is incredibly helpful. Yeah, and as an example, we talk about test doubles, and broadly in the industry everyone just talks about “mocks” in general, but there are semantically different flavors of the tool. And I think Tom Manshreck — who’s also on the call here — tracked the terminology from Martin Fowler originally, like 20-odd years ago, and we sort of reconstituted a lot of those ideas. Once you start explaining the different use cases — you could just have a flow chart that tells you how to choose what the appropriate tool is for this task — so much muddle just crystallizes, and it’s kind of magic to watch.

George: Actually, I can speak to some of the testing stuff, because I can attest to Titus’s part. On the testing side, Titus hammers vocabulary about flakiness vs. brittleness — which for many folks is just a pejorative adjective to slap on a bad test. And he’s like, no, they’re two distinct things with two distinct remedies. So it’s important we give these names. One thing Titus may not be aware of: I’ve been fighting the fight to use the term “quality attributes” instead of “non-functional requirements,” and I think we might actually be over the hump on that. Believe it or not, having the website up for five years, people are actually starting to say it in authoritative, very senior-level documents. So thank you, Leah Rivers, and a bunch of other folks — but I think we’re actually winning that one. Another answer: there’s almost no Google technology in any of these topics. You’re not going to see Stubby, you’re not going to see our internal database names, any of that kind of stuff. It does lean toward using examples from languages we use at Google — not Haskell or ML. So mostly it’s conceptual.

David Garlan (audience): So what are the prospects that this might become available outside Google?

George: We’re working on it. It seems entirely possible — that’s probably all I can really commit to at this point. There are a lot of people who’d like to see it happen. Because I think you and I have chatted about this idea I’ve got, which is: it’s preposterous that this is a single source, that I’m the only one contributing. I would love a Wikipedia kind of thing where everyone can start to improve this thing. I know Wikipedia sounds scary, but over time, with editors and attention, it has become truly impressive. I don’t see why we couldn’t do something similar. Or you could have it be a CMU-branded one, or you could have different forks and say, well, that’s the Stanford flavor of it, but here’s the CMU flavor — like the Linux kernel. Anybody can do anything you want, but we tend to follow Linus’s version.

Audience: Yes. One of the things you mentioned toward the end of the talk: you see a lot of code in industry that’s like 1970s code. The implication is that there are characteristics of this code that make it poor quality, and you mentioned a lack of use of abstract data types. Can you describe more of the characteristics or deficiencies of some of this code, to give a better idea of the challenges facing the industry in code quality?

George: Yeah. So there’s a thing that’s going on in industry, and has been for a while, which is that the majority of the code is written by the least experienced programmers. The more senior you get, the more you get into an advisory, steering capacity, and you don’t get to write as much code. I think that’s less true at Google than at many other companies, but it’s still true to some extent. You have an army — it’s a business model: one expensive, experienced person plus five other people to magnify their influence. That’s what you want. So you end up with mundane rookie mistakes being magnified. An example: engineers across the industry, I think, are not properly afraid of side effects. When you’ve got a hundred lines of code, it’s easy to keep track of the side effects in your head. But when you have tens of thousands or millions of lines, it becomes very difficult to reason about “I just made this change — is that going to be safe?” if you don’t have some discipline about where things are going to change and where they’re not going to change. This is an example of me personally speaking for myself: it took me a long time to understand many of the things the programming-language community was saying. Now I can see they’re doing languages in a certain way, but I’m not allowed to use that language at work, or I’m not using it — the system’s already written in Java or whatever. So I’ve been doing my best to mine all the good ideas I can find, but state them independently of the language that carries them. So here’s a page on why you should be worried about side effects: they don’t compose very nicely. First of all, the idea of composition of code isn’t something entry-level programmers really think about. And the idea of the contractual nature of code — “I give you an X and you give me a Y.” Even in procedural code, if something is non-contractual, it quickly becomes a kitchen sink: if there’s no contract, then when I have to put another feature in, I’ll just put it right here. So the better they can hit the target of “this method has a clear purpose,” the less likely everyone who edits it is to start adding weird stuff. Does that give you a flavor?

Audience: Yeah, it does. Thank you.

Audience: Hi, George. You mentioned that there are both go links that the SWEs can reference for the concepts, and there’s also Gemini that adjusts based on the links provided. Do SWEs usually consult Gemini more, or the go links more? Under what circumstance do they lean more toward consulting Gemini as opposed to the go links?

George: Well, unfortunately, the answer is it’s too early to tell. The use of LLMs in industry is just beginning. For a long time, the LLMs were only looking at external content. So me seeing them look at internal content, I was like, “yay, it’s finding internal stuff.” I expect this trend will continue, and I’ll be able to answer your question a bit better later on. But the one distinction is that in this kind of scenario, your tech lead is specifically reviewing something that you are doing.

Audience: And I would imagine LLMs might often be used prior to sending it for review. I think LLMs could also help with reviews.

George: But in this case, this is someone else trying to communicate with you via a compact essay, if you think of it that way. This is the essay they would have written if they had the time.

Audience: Thank you.

Bill Scherlis (audience): So thanks for your talk, George. I’m curious about follow-up. The TLs go through the class, go off, do things, and then they answer your questions, and they like it. The question, though, is: after maybe four months, six months, they may have some issues that show up in their experience that are late-breaking — not anticipated in your instruction or in their early adoption. Do you have a mechanism for refresh, recharge, revisit — do the TLs come back after some period of time and say, “great class, great ideas, except this one thing, we just couldn’t get any traction on it”?

George: So this is a part I want to point to. We’ve got the written materials and the courses. We have this podcast kind of thing where we invite the alumni to be the guests. I assume the same thing kind of works with Zoom — with Google Meet anyway. You can be invited to the meeting, or you can have a read-only link and watch the live stream. We invite the whole company to the live stream, and the alumni get to be in the meeting so they can raise their hands. So there are certain things we’re doing to try to provide a relief valve for that. In the early days, we made a conscious effort to forge a community around each one of these cohorts, because in some areas that works well — like I understand MBA cohorts really stick. My brother got an MBA, and he stays in touch with those people as he moves forward. We found we could do that, but the effort on our organizing side was very high — we had to organize lunches, send emails, and do various things to keep it alive, because it’s just going to decay back to nothing. So we were unable to keep up that level of effort; we didn’t have the staffing to do that. The idea of having them actually come back — I love that idea. I love that idea. Yeah, like a follow-up. Help us. Yeah, they tell us what they don’t like.

Audience: So with organizational change being a peer-based hearts-and-minds game, I’m curious: were there any commonalities among the tech leads that determined successive mentoring, or successive adoption of the techniques?

George: Well, I’m trying to think back on what the feedback said. One of the things they said in the feedback was that when one person went back to their larger group, they felt like a voice alone in the forest. But as soon as they could advocate to their management to get two or more people, it started to get traction inside the group. So they keep coming back to this idea of critical mass. If you have three tech leads saying “yeah, I think this is a good idea,” it stops feeling like George’s screwball idea and starts becoming “yeah, a lot of the people are talking about this idea.” But you’re exactly right as far as a social dimension, the organizational-change dimension, which I’m a novice at those topics. I’m sort of wandering into it and doing my best.

David Garlan (host): So, George, I’m sure there are many more questions people could ask, but we should stop. Let’s thank George again.


Transcript prepared from two ASR passes of the talk audio, cross-checked against the recording. Timestamps are approximate (±30s).