Saturday, July 25, 2026

Why learn theory


Why learn theory? We were bringing our daughter back from an arts camp in northern Michigan. As we talked about music, her lessons and teachers, pieces, performances, and ensembles, we also talked about her music practice (music theory and music history) class. And as a musician, while music practice was a welcome and sometimes entertaining break, she asked why we learn theory.

All practical domains have an underlying theory. Theory is how a domain understands its environment and how practitioners and researchers interact with its environment. For example, knowing the fire triangle/tetrahedron enables firefighters approaching a scene to assess how best to control and contain a fire while ensuring the safety of people in the area. Planning looks a capability and capacity and how to employ these to turn strategic and operational goals into resource allocation and actions over time.

What does theory give you? First, it helps explain why things work or do not work. Theory provides the basis to understand success and learn from failures. So that each observation is not simply a succeed/fail assessment, but you learn lessons on what made it work and how it can be done better. And each event/observation is also an opportunity to examine the theory and make it more detailed (with experience and wisdom comes a more nuanced and detailed theory. e.g. physics)

What are alternatives to theory (always know the alternative to any concept)? The most common is 'how we have always done it" or "this is our experience". Next is "loudest person in the room" and its relative "highest paid person in the room". But experience is highly situational, and without theory you do not know what is the key detail that made the experience what it was. The appeal to authority and judgement is also highly situational - people's expertise is in large part based on the experiences they have had (one reason that best practices for developing executives and general/flag officers is to give them experiences in a wide range of settings over the course of their early careers). Theory creates a framework to have conversations and identify what the questions that need answering and how to create and evaluate alternative solutions.

In my own area of operations research, which has led me to artificial intelligence and it generative AI form, theory is what I use to evaluate alternative approaches to addressing a problem. I like to say I know how to break every tool I've ever used. In this new era of Generative AI, I start projects with new business partners and domains by figuring out what about the problem domain generative AI gets wrong. (my business partners say that I am secretly happy when they find mistakes or cause problems with generative AI tools. They are wrong. I'm openly happy) And theory helps look at those problems and work out ways to modify those tools and make them work in the setting at hand. Even as the world changes around me.

Monday, February 09, 2026

Claude's Constitution on the role of the analyst

As I was reading Claude's Constitution, I was reminded of discussions on what the role of an analyst advising decision makers should be. One of the core beliefs is that the principal (the decision maker we were advising) needed frank advice on the topic at hand so that they can make informed decisions. What makes this hard is a strong pressure to either do what the principal asks for or to say what the principal wants to hear. The Claude Constitution (and generations of analysts in the liberal democracies) reject this and characterize its role as genuinely helpful, courageous honesty, and a commitment to the organization's long-term well being. Part of socialization of being an analyst in the United States is valuing frankness and competence. And this is taught person to person, through example and stories of bosses who wanted analysts to only do what they are told, analysts who worked to give their bosses what they wanted to hear, and the consequences of bad decisions that hurt the organization and mission, even if the analyst and boss felt good in the moment. It is a story that is retold in many novels and movies as plot points leading to a disaster. And when military and intelligence analysts move to the private sector, one culture shocks is an American management culture that values obsequious servants instead of frank and competent advisors. In this constitution, Anthropic is firmly on the side of frank and helpful, even if it causes discomfort for the user. This is in stark contrast to the common observation that generative AI is eager to please and wants to tell users what they want to hear to maximize engagement. One thing about the constitution, Anthropic wrote it knowing it would not have an opportunity for conversations over time, which is how values like this are taught. And since this is written for an AI, they could be verbose. So there are a lot of different aspects of the question of how to be a good advisor that are discussed, each a slightly different take on what it means to be genuinely helpful with a myriad of nuances. In particular, it has many discussions along the lines of "what if the user wants only affirming answers" and "what if the user does not ask to not be given misleading statements that will lead them into trouble". Because that is something that is seen in real life. If I were a place that taught analysts, like schools of policy or the service graduate schools, I would assign the Claude's Constitution as a reading assignment. And there are a lot of topics that are covered here that would be topics of reflection and rich discussion. Because these are issues that analysts will face in real life. And history is full of examples where obsequiousness leads to harm for the organization. And those who claim that mission and service come before self need to reflect and discuss this in the classroom, so that they are prepared for these issues when they occur in real life.

Tuesday, February 03, 2026

Book review: Little Brother by Cory Doctorow

This was written in 2008, in a U.S. where a Department of Homeland Security has turned on Americans in the name of security following a terrorist attack. And as its focus and energy was on keeping Americans in line, it created opposition in the form of people who wanted to go about their lives and do so without surveillance. And in this, the protagonists are teenager who were caught up in a dragnet while playing games. I was in a group that had a conversation about computer privacy recently, and someone expressed a belief that everyone in positions of power was supportive of encryption and privacy. But I remember the days when governments actively opposed the spread of good encryption, even for businesses who wanted to be able to secure discussions and transfer of payments. The technology issues brought up here I still remember, and I recognize a lot of the debates and issues, the opposition to encryption, the desire to do mass tracking and surveillance through cameras and electronic devices. And the technology and social methods that the kids in the book used I recognize from those that were used in real life (which fits the author's purpose). But the real question here is what does being an American mean? One scene involved the review of the Declaration of Independence and what it means, all the while comparing the grievances expressed in the Declaration of Independence to the fictionalized Department of Homeland Security. The kids are not a White Rose brought to present day. They are not philosophical about morality and righteousness. They only want to live their lives as they see fit, and for their friends to be able to do the same, because they recognize that if the right to live as they please is refused to some, it is refused to all. So they act in ways that provide the freedom action they want to everyone. In contrast, the purpose of those who are demonstrating power is purely to be able to exercise power and limit the ability of others to live their lives. Another side note. Power for its own sake becomes incompetent. Free creativity creates its own quality. The kids in their creativity create technical solutions outpacing the Department of Homeland Security, who only knows power and cruelty. The kids are not super technologists. They peripherally encounter older protagonists who can do what the kids do, but better. And because they are focused, the older counterparts are quieter and more effective. But this story is about the kids. This is not a tech-savvy version of White Rose. The kids are not portrayed as idealistic saints. They are kids, who want to live freely and create space for others to do so as well. Which is a very real way what the Declaration of Independence proclaims.

Sunday, February 01, 2026

What is needed to work in the age of Generative AI

Last week I was at CMU-Heinz for a fireside chat type event with students in the various MS in Analytics programs there. One question that I got was what were the skills needed to succeed in an environment with AI, and even into the future.  Then I spoke about being able to program because you need to learn how to think deliberately, being able to connect technical capabilities with end business needs (because this has always been how analytics fails), and as I think about it more having a better understanding of knowledge. Because if you believe that your education and training is about learning sets of facts and recipes, AI will eat you alive. So your understanding of your field has to be greater than facts and procedures.


First, why learn computer programming when AI can write code faster than you. Microsoft has a set of studies showing a high (40%) rate of errors, yet their programmers also say they are more productive. Because, as I use AI at work, I find that it is helpful in creating good structure and framework scaffolding, especially when I have to context switch (I regularly switch between three data stacks at work, each of them have many people who spend all their time in one) or if I am applying methodologies new to me or my organization.  But because I am competent, I can correct it as I go, and the fact that there were originally errors is not a big concern, because I was going to revise everything anyway


Another reason to learn programming is you learn to think in a different way. The ancient Greek philosophers had students learn geometry before philosophy. Not because geometry and math is beautiful (even though they are), but because with geometry comes proofs. And geometric proofs is about how much you can understand starting with a minimum amount of assumptions (Euclid's five axioms). And you now have experience in determining an objective truth, no appeals to authority, no claims of different point of view. And your logic is in the open, to be critiqued on their own merits. Far different from my friend in grad school who claimed that perception is reality. And only then were you fit to move into the realm of ideas, where even facts have to be evaluated.


Programming languages differ from natural languages in their precision. Every statement has a single clear meaning. And this is different than natural languages, where the ambiguity of human life plays a role. So to work with anything regarding computers, it is helpful to recognize that computers will work with language in different ways we do, it handles ambiguity differently than people, and how it will use randomness to handle the difference (which is key to how Generative AI works).


The next is the link between the capabilities of technology and the needs of the business.  According to everyone who has studied project failures in depth, failed communications between the business partner and the analysts is the biggest cause of project failure. And data projects have a failure rate between 80-90% (this range has been persistent in studies over decades in the data world, and it is consistent across definitions of failure and different segments in data analytics, data engineering, or data reporting (dashboards)). Being able to understand the business needs of end customers as well as understanding potential classes of technology solutions leads to asking better questions and getting value out of the applications of technology.  The main reason for breakdowns in communication is ego and arrogance.  From the technology side, there is often a belief that the customers are idiots who do not know what they want, so the technology people should just build something and pitch it back to the customers.  This is mirrored by business people who think that technology is a a turnkey product so they should not interact with the people who are creating the solution.  A third variation is when upper leadership decides to act as an intermediary between the analysts and the end customer. The logic here is generally that the leader believes both the technologists/analysts and the end customer have no communications skills, therefore the leader will handle all of the communications and give the requirements to the analysts.  All of these are wrong. Especially in anything involving data, details matter, and the entire project involves discovery of details that no-one realized were important at the beginning. So the analyst and the end customer need to be regularly reviewing these discoveries, and adapting along the way.  And only the end customer (because they are closest to the problem on the ground and know what kinds of actions can be taken) and the analyst (because they will be representing the detail in models and they know what the range of alternative models can do) together can make those decisions.  Without that direct communication (potentially facilitated by someone who knows both sides), a project falls into the trap of solving the wrong problem. And this requires people who understand purpose and can determine the impact of nuance, both of which Generative AI does badly in.


The third category of future work is understanding your field.  Computers are very good at retrieving facts, if those facts are in its knowledge base. Gen AI is better than prior technologies because it is not as sensitive to getting the wording precise.  Computers are also very good at following instructions if those instructions are given. (people also tend to be better at things when they are given good instructions). So, if this is the extent of your subject expertise, you are in trouble.  In software development, there is actually a very large workforce like this, whose careers are built on the ability to fill out a given framework or instructions. But if your place in the world is built on more than knowing facts or following recipes, if there is actual understanding that has to be applied on a situation specific basis, there is still room for you. Without that understanding, an organization can execute perfectly, a solution to the wrong problem. Which is worthless. So you need the level of understanding that allows for good judgement, and you need to be working in an organization that allows for its employees to use that judgement.


A last criteria is based on a number of conversations I've had.  Many people express that they believe in the answers Generative AI gives because of the massive investment these companies have made, the smart people they have hired, and the belief that these companies would ensure correctness. I had to explain that these are industries and communities that have historically claimed they had no interest in accuracy or correctness. And until recently, viewed the paying customer as king and only sought to full the demand. Ethics was not part of the conversation.  And what they delivered, did not come with guarantees other than it does what it does.  This attitude that you did not question the authority that came with wealth and success is the first thing that has to be broken before people can use Gen AI productively.  Both of my kids do it.  I design the rollout and presentation of projects at work to make sure my business partners who are using Gen AI view its output skeptically and looking for specific types of flaws.  As an avid reader of science fiction over the years, much of which addresses AI as part of society, and worry much less about the power of AI than I do about people who use the output of AI without being critical thinking. It is the kind of following that leads people to enact policies without analysis, and punishes people. And the outcomes are the fault of the people who followed the AI. Because AI has no goals, purpose, or conscience beyond that of its user.

Wednesday, January 21, 2026

Book review: Tools and Weapons: The Promise and the Peril of the Digital Age by Brad Smith

 

Tools and Weapons: The Promise and the Peril of the Digital AgeTools and Weapons: The Promise and the Peril of the Digital Age by Brad Smith
My rating: 4 of 5 stars

The author, Brad Smith, was General Counsel for Microsoft (he was also President of Microsoft, but his role as counsel is more relevant for this book). The book is a discussion of privacy in a context where governments and large corporations hold immense amounts of personal and business data, and there is a large temptation for corporations to take advantage of that knowledge or governments to access that information, for governments in pursuit of legal action or suppression. So much of the book is about cases that involved Microsoft and how they developed a stance on corporate responsibilities to their customers on privacy matters, specifically in cases where government demanded customer data.

Each chapter revolves around a policy argument that played out in public forums, regulatory, legislative, and in the courts in the U.S. and Europe. And it is in a backdrop where technology companies used to believe that as technology companies they did not have an interest in policy. But as companies became less sellers of goods and more providers of services, in particular of data storage and cloud based communications services, they became targets of government and criminal action to access customer data without consent. In each chapter Smith introduces the context, then introduces a historical principal that predated cloud computing, and he makes the argument that the choices and policies used to govern oud based computing services should be the same that governed the same type of services in the pre-digital age.

The overall philosophy he gives is that Microsoft is a custodian of customer's data, not the owner. And as custodian it will protect the customer's property (data). And throughout the book he identifies allies (who have similar philosophies of protecting customer's/citizen's property and privacy) who only differ in details. And those that he as to be contentious with, because they are seeking to use and profit from individuals data or desire access for investigations. (and Microsoft in these cases wants a transparent process for doing this what protects their customers, who have the rights of citizens/residents)

Clearly, Smith is proud off his work, and believes that protecting the privacy of Microsoft customers, even in the face of government pressure, is the right thing to do (with a procedure for governments to prevent harm to other citizen's rights, life, or property. But he does acknowledge allies, Google and several European governments come across very well here. So a reader has to be mindful that he does have rose colored glasses on Microsoft's journey in this topic.

I appreciate the view of a non-technology person on these topics. As he is a lawyer, his perspective is to look at issues that seem very new because of the pace of technology change, and recognize that the issues have existed and debated before the digital age. As the infrastructure is owned by multi-national corporations, the relative power of industry and government is different. But the idea that industry desiring to protect the interests of their customers and government desiring the safety of its citizens should align is one worth engaging in.

View all my reviews

Monday, December 29, 2025

Reflections on the Advent of OR: Using Generative AI in Analytics and Agile Operations Research

In December 2025 I participated in the Advent of OR (https://adventofor.com) which was a 24 day exercise that guided participants through an optimization project. And instead of just solving problems and creating models, the Advent of OR walked through an entire project life cycle, using the INFORMS Analytics Framework.


While I am not a student or early career who was the target audience, I took part, and I had three goals.


1. Use a new programming toolkit. I used VS Code with R and Quarto.  I usually use R Studio and I wanted to try R on VS Code.  And I think Quarto is the future replacing Jupyter Notebooks for Python and a natural evolution from R Markdown.

2. Practice in optimization. In the Operations Research world,  I am NOT an optimization person. My thesis was applied probability (queuing) and my methods research has been in simulation (one stream in ranking & selection and another stream in Bayesian methods for input modeling)

3. Using Generative AI. I wanted to see how generative AI does in an operations research project.  And I wanted to do it right in a setting where I can give it references to guide it.  Note: I have found that Generative AI favors descriptive statistics, machine learning, and hypothesis based statistics over other forms of analytics, so it needs some guidance.


Toolkit


I had to set up VS Code with the R extensions, Quarto (and extension), ompr and the glpk with associated R ROI packages, and to make sure everything worked, I download the repository for OR_using_R by Tim Anderson.  Then to render the book (meaning I made sure all the code ran) I had to install texlive with xetex and extra fonts.  Generative AI (I had Gemini CLI installed) was very helpful in all of the system administration tasks since it could figure out what was needed every time there was an error message.


Data analysis and Optimization


Working with the data sets, it read in the data, (I had to give it some corrections along the way to help it recognize the data types.  When the data files were read in, it recognized that the data sets did not correspond in granularity.  In the R markdown file it created, in addition to generating the code that read in the data and created summaries, it also identified a number of questions and concerns about the data and created questions for the stakeholder.  This was a good set of questions that corresponded to what others put forward.  


It also did well with the optimization.  Given an optimization textbook, I first asked the generative AI for a mathematical formulation based on the project description. Similarly, it created a process for determining what kind of problem this was and worked through that process to determine that this was a linear programming optimization problem.


Next was a LP formulation using OMPR.  The first formulation was straight forward.  I went and had the GenAI break out the formulation into its own R script to enforce a separation of concerns between the data handling, optimization model, and output processing.


I generally also ask for docstrings as I go, and the Gen AI did this for both the model as well as various handling functions. I generally read the docstrings to ensure they say what I expected them to say. When it did not, since the docstrings were written based on the code, I took it to mean that the code was not right (this exposed a mistake in the initial formulation of the LP in OMPR).  Similarly, I had the Gen AI write unit tests for the constraints and a mock problem to test the optimization.


Agile Operations Research


One of the aspects of having the Advent of OR over 24 days is that it rotates topics between the art of modeling, implementing and managing models, and interactions with stakeholders.  There are a couple of very important points. First is that interacting with stakeholders is not something that is at the beginning and end of project and ignored in the middle.  There needs to be stakeholder engagement throughout the modeling process.  A second point that has come out in the conversations on LinkedIn is that the most common cause of project failure across data analytics are communication failures, in particular between the analyst and the end customer. While this can have many causes (including management inserting themselves in between the analyst and end customer), as analysts we must have that direct interaction from the beginning of the project (business problem formulation in the INFORMS Analytics Framework)


In the early stages of the project, one factor we need to face is failure of imagination. The first level for analytics is that our stakeholders often do not know what is possible across the full range of analytics.  Often a problem is presented as a request for a tool, but for the results of the project to have any value, it has to address the end problem, so business problem formulation has to start with the end problem, determine what kind of information from data can help the decision makers address the problem, and then we can start discussing what methods can provide results in the form that will be useful.  Currently, because of media hype, the initial request can be for a dashboard, or a predictive model, or a generative AI tool.  As operations research analysts we can also bring to bear statistics, forecasting, optimization, simulation, and queueing; and different ways of applying those methods to give different kinds of results that can be delivered to decision makers to make better decisions.


After the business problem formulation, the next big change in the project will occur when presenting the first minimum viable model to the end user. This is the first model that uses a minimum acceptable subset of the data and model that covers the most essential aspects of the smallest version of the problem. The reason this is important is before this, all conversations are abstract and theoretical. The first time a model with outputs is presented to an end user, the end user will start to imagine how they would use these results in real situations that have happened in the past. And they will start telling about all of the considerations they account for, the information they need to gather to make decisions, and who they need to consult and coordinate with. And this can change the entire project.  And from experience, I do not think it matters how much work is done at higher levels to define the project, the first time a model is presented to an end user the project will change so that the outcomes can be usable to the business. So it is best to make that happen as early as possible so that change causes the least disruption to the work in progress.


The idea of rapid cycles of iteration and feedback from the customer, and the willingness to accept changes to the project due to that interaction are the hallmarks of agile development methodologies in the software development world. Having regular rounds of model iteration where additional elements are added to the model, and getting feedback from stakeholders to confirm that the project is on the right track to produce something useful. And just like the software development world has experienced, this is more likely to lead to useful product, and actually faster than attempting to follow a rigid path that leads to something irrelevant.  


Conclusion


The Advent of OR proved to be a valuable exercise, offering a full-cycle project experience that highlighted two critical modern aspects of Operations Research: the integration of Generative AI and the necessity of an Agile approach. Generative AI demonstrated significant utility in accelerating system setup and basic modeling tasks, freeing up the analyst for higher-level problem-solving. More importantly, the experience reinforced that project success hinges on continuous, direct stakeholder engagement, mirroring the principles of Agile development. By prioritizing early delivery of a Minimum Viable Model, analysts can gain crucial feedback that aligns the project with real business needs, ultimately reducing the risk of communication-based failure and ensuring the final product is relevant and utilized.


Monday, November 17, 2025

Thoughts on mentoring within the analytics profession

Our careers and lives follow unique trajectories and structures. While we all have our own paths, it is helpful to have people who have gone ahead on similar paths to share experiences and thoughts on the future. Part of our professional development are mentorship relationships, which can be done in a wide range of settings, relationships, and time frames.

I am going to define mentoring as a longer term, unstructured professional relationship, with the focus of the relationship being the personal growth of the mentee.  Typically, the basis of the relationship is that the mentor has gone on a path that the mentee is on themself, and the insights of time may be helpful for the mentee's development.

One thing that distinguishes mentorship relationships from other professional relationships is that mentorship relationships are holistic.  They look more than just the task at hand, or even a job position. The mentorship relationship may be career focused, but it will look at the whole person, and will recognize that overarching goals can change with life events, even life events outside their occupation. So, while a supervisor/manager can be a mentor, this is really not apparent until after the manager relationship has ended, and the relationship has become larger than the roles both individuals had when the relationship started.

As we all have unique life paths, we cannot expect that any one person has gone on the same path that we are on, but mentors bring not only their own life experience, but also the experiences of those whom they have lived life alongside. They have seen the decisions and choices of others, and how those decisions have advanced the goals, or not. They have seen people whose lives have taken them on different paths, and so have a broader view on what the future can hold than those whose view of the world is from the relatively structured life of home and school.

What topics come up? The focus on a mentorship relationship is on the growth of the mentee. In the context of technical professionals, this is the professional growth, but as part of a full life. So, with an understanding of the long term goals of the mentee, it can be working through broader issues on a project, such as other points of view.  It can be soft skills or relational skills working with co-workers, superiors, juniors, or outside colleagues (customers, business partners, etc.).  It can be suggestions on how to stretch as a person, to see and work through things from a broader perspective, and the skills needed to do this.  A mentor can be a sounding board, providing different points of view (especially on the behalf of people who may not be good at communicate their point of view). It can be how to handle work/life balance, looking at a whole person. It could also include looking at alternative paths, that different positions or even career paths may be more suited for the goals of the mentee. 

How does a mentorship relationship start? Like all relationships, you can never tell if a relationship is going to be long term at the beginning.  But you have to begin somewhere. A first conversation is often about a particular topic, one that is of mutual interest. (and this initial meeting is sometimes arranged by organizations such as a company or a professional organization trying to promote mentorship among employees or members). After the first few conversations about that first topic, you should have observed if the relationship is broader than that first topic, and you can talk about if you want to continue meeting about topics as they come up.

What does the mentor get out of this relationship? Typically, people who are in mentoring relationships also have other rich relationships, which is how they get the background that makes them valuable as a mentor.  Over time, the relationship becomes driven by both concern and curiosity about the other's experiences in life. Often that includes issues that are more apparent to someone at an earlier stage of life or career. A mentorship relationship can then become one of an ongoing set of relationships that makes up a life well lived, and the ultimate hope, even when it is not an expectation, is that a relationship be one that lasts.

Do mentorship relationships last?  Sometimes. Organizations such as workplaces and professional societies will often organize mentorship relationships, but these are always based on a topic of interest in the moment, and these relationships typically start out with short term boundaries. But, like all relationships, a short term relationship is what has potential to broaden into something longer. Does the relationship broaden beyond the topic where it was started?  Do conversations evolve organically and feel natural when they branch into new topics? Over time, can the relationship feel like something that lasts as both sides grow and change (as all growing people do). So the transition from a formal, temporary relationship with a defined schedule and defined boundaries changes into something more long term and fluid. And a mentor/mentee relationships begins to feel more like professional colleagues, each moving through life and careers on adjacent paths. 

Can mentorships relationships be informal? Yes, in the sense friendships are informal. In professional society meetings, it is common to see someone and immediately follow up from a conversation from a year ago, just like old friends. So you can have a relationship where you only see each other on occasion, but immediately pick up where you left off, just as old friends do. But the key is the long term relationship, that the conversations are about growing people, not only about topic at hand.

Are there aspects of Analytics that mentorship relationships are especially helpful? One area are the soft skills, the skills of working with colleagues, managers, and customers that is not part of the standard training of a technical professional. A mentor can relate to what the other person may be thinking and help the mentee develop that sense of empathy for others that make them more effective professionally.  A second aspect is dealing with the hype that often accompanies the profession. The most recent example is the rise of Generative AI, but similar waves of publicity occurred around deep learning, big data, and machine learning in general. A mentor can place new ideas and concepts in the context of everything else a mentee knows, in contrast to teachers or thought leaders whose responsibility at any given point in time is single topic focused. A third aspect is a sense of what a mentee may need to be a well rounded professional. Training programs and classes tend to be singularly focused with a specific goal, but professional growth needs to be holistic, and designing such a path needs the attention of a person who is looking at the whole person.

Mentorship presents the potential of a valuable relationship, fostering personal and professional growth through a holistic and potentially long-term connection. It goes beyond task-oriented guidance, embracing the mentee's whole person, from developing crucial soft skills and navigating career paths to contextualizing industry trends. While it can begin focused on specific topics or within formal programs, at their best mentorships evolve into enduring relationships, offering mutual benefits and enriching the lives of both mentor and mentee. In dynamic fields like Analytics, such relationships are particularly vital, providing the comprehensive support needed to cultivate well-rounded, effective professionals in changing times.

If you are interested in mentoring relationships, I would look to your professional society. If you are in analytics, I would recommend you look at INFORMS and their mentoring programs (Video on the value of mentoring in analytics)  It is a professional society for advanced analytics (broadly defined) and is vendor, tool, and methodology neutral, which is important for a field that sees major changes over the course of decades.

Wednesday, October 22, 2025

A tale of two Corne: one month with split keyboards

My split keyboard journey started with two Corne keyboards and a Sofle all purchased over a period of two months. The Sofle is from Ergomech and is a Bluetooth enabled.  I use it as a wired work keyboard.  The two Corne keyboards are from YMDK, purchased through Amazon. One is an MX, the other is an MX low profile.  Here, I talk about the two Cornes. I will look at (1) Buying from YMDK on Amazon, (2) use of the two keyboards, (3) the keymap journey.






I bought the keyboards from YMDK on Amazon.com. YMDK markets pre-soldiered, hot-swappable,wired and wireless  (2..4GHz dongle) Corne 4.1 keyboards with 3D printed enclosed case with 46 keys  (3x6 +5). Also a wireless Sofle.  I wanted wired only to make my first steps into the split keyboards with fewer complications, in particular the wireless versions are powered by replacable button cell batteries, that I did not want to deal with. I immediately flashed both keyboards with 4.1 Vial versions of the Corne firmware and that had no problems.

I have one keyboard with MX Akko Dracula linear switches (35 g weight) and XDA profile PBT keycaps, and one low profile with kaith Deep Sea Whale Low Profile choc v2 (silent tactile).  The first keyboard I ordered was a refurbished MX switches. The right side did not work, and I suspect this is because Amazon refurbished items usually are returns, and the prior owner probably shorted the keyboard. A couple back and forths with YMDK customer service (there is a link on Amazon to them, and it is not too hard to get the YMDK customer service email address.) and we decided to return it through Amazon and I ordered a new one. The Low profile keyboard was not a problem. The website make it clear that it could take either Kaith 1333 (v1) or 1353  (v2) switches. So I got 1353 switches and used Wormier Low profile (skyline) key caps.  So, with Amazon return policies, I found buying from YMDK and working with their customer service reasonably good, although I am leary about any complications like wireless.

For using the keyboards, I use them with my own laptop and I have one at a standing desk that I use for both my work and personal laptop (I take a standing session and connect my laptop to a docking adapter). And I also bring the low profile Corne with me when I go in to the office (my acrylic sandwich Sofle looks a little fragile with its openings and the suspended acrylic oled screen cover). I like the low profile version. With the 3D printed case and low profile keys and switches, it feels relatively durable and no big failure points like catching on something. I warp up the halves in a bandana and put it with the cables into a lined bag and that seems to work well. I'm not sure I would like the low profile keyboard as my only keyboard, it feels tough because I'm always bottoming out compared to my MX Sofle. But as a secondary keyboard to provide variety I think it does well. And it does get attention when I take it around :-)

With the Akko Dracula switches, I think the light switches don't work for me. I am constantly having accidental key presses with mod-tap keys that I don't with the silent tactile keys with 45g weights. I think I'm going to put that keyboard aside until I feel like getting new switches for it.

The keymap journey is an ongoing one. I think I'm at a point that I only make small changes a few days apart.  Some big choices along the way in roughly the order I settled on them.

  • QWERTY- I'm staying with the QWERTY layout. I know it well, and I am not so fast a typist that any keymap optimization can make any meaningful difference.


  • Numbers. I started with numbers being a top row of a layer. Eventually I realized that when I need numbers, I need several, so I switched to making a numeric keypad with arithmetic symbols on one side, then the other symbols on the other side of the layer.
  • Symbols. There are seven'ish pairs that need to be taken care of. I touch type, so I wanted the pairs that are usually on the same key together.  So these are:  `~, -_, =+, [], {}, '", \|.  The quotes get taken care of by moving enter to the thumb cluster, so quotes stay in place. =+ and -_ I put with the numeric keypad. So I had two rows of the symbol layer were the []\'` on one row and the shifted version of those keys {}|"~ were in the row below that.  I put the brackets on the outside (left row) because I ended up putting all the brackets on combos, so I still have them on this layer, but on the edge.  And since I program in R, I put the ` and ~ closer to the index finger.  The top row of the symbol layer I put logical operator symbols:  <>&|!, (less than, greater than, and, or, not).  Making two columns of all the bracket keys (except parens) and keeping them in an order I would be able to remember.


  • Navigation. The top row of the navigation layer are the symbols that are the shifted number row.  For the remaining two rows, the right side is navigation, left side mouse control. Right side is centered on hjkl, because I use VIM and my fingers already know those keys.  Below those are horizontal keys, beginning of line, previous word, next word, end of line. I put page up and page down on the two keys to the right of those. On the far right I have beginning and end of document, but I don't think I use those that much. The mouse cluster is xdcv.  Then f,s are the click buttons, gb are scroll up and down.  za are scroll left and right.  I use these a surprising amount of the time (especially to help recover after accidental mod presses)


  • Adjust layer:  Left half is for controlling the keyboard (lighting), right half is media controls.  I have volume mute, down, using hjk, media back, play/pause, next anm,  l; are screen brightness.  ./ are zoom out/in.  I never really learned to use the keys because these were always on the function row and every keyboard had them in a different place. Now that I got to put them where I wanted, I use them a surprising amount.


  • Screen navigation.  I had window management (switch windows, move windows) on the thumb cluster. I figured if I was in this layer I did not need space, enter, backspace/delete. I knew these keys, but they were awkward on a normal keyboard
  • Programming key combos.  I made combos for the bracket symbols ()<>{}{} with left on the left side and right bracket on the right. I tend to use these instead of the normal typewriter layout.  I have additional combos for : <- |> # for R programming (the I also have combos for open file and the command palette in Visual Studio Code.,
  • Mod-tap and layer-tap. I have an extra layer tap keys on g and h, which I use to mirror my layer keys so each layer is accessible from each side of the keyboard. For example, the number pad, I can either use it single handed with the thumb holding the layer key, or use the index finger on the other side. I usually use it one handed if I only need one key on the number pad, opposite hand if I need more. I let my fingers decide. I also have (from outside in) home row mods of layer, shift, control, alt, gui. But I took out the gui modtap as I was doing too many accidental mod presses, especially with the Akko Dracula switches. (I think the other keys are not as noticable because they have momentary effects, but GUI and menu lead to something happenning.)
  • Thumb clusters, I have ended up with the thumb cluster being Insert(held Ctrl), Enter (held GUI), raise layer (navigation), lower layer (numpad/symbol), space, backspace (held Alt). I realized that the menu key was also right click on the mouse (and the key on the mouse layer works), 
Some other decisions along the way
  • Delete/Backspace. I first left backpace next to P and delete key in the thumb cluster, but I repeated used delte when I meant backspace, so I moved backspace to the thumb cluster and had delete in the corner.  It makes for the same Ctrl-Alt-Delete chord that I'm used to.
  • Escape and tab. I started with escape in the corner, but my fingers always wanted tab to be next to Q. So Tab went into the corner and Escape went where CAPS LOCK usually is. I don't use CAPS LOCK much,  so I made a combo with both space keys as something easy to remember if I ever want it.
  • Numpad. I started with the  numbers on the top row of a layer, but they were always awkward, just like they are normally. Then I realized they could be a numpad, with room for arithmetic keys around them. I also tried both left and right sides, and ended up with the right side. Because - and _ are used so much, I had those under the index finger and =+ went to the other side of the numpad.
  • Shift and Ctrl/Alt keys. I tried out putting Shift in the thumb cluster and Ctrl, Alt in the corners where shift usually was, but changing that muscle memory was too hard.
  • Space and Enter. I tried space on the left first, then saw I was making too many errors so I switched them.
Observations from use.
  • I don't miss the number row. The only time I notice it is when typing passwords or other things like phone numbers where muscle memory knew where the numbers are, but I'm creating a new set of muscle memory.  And the symbol keys always needed a layer key (Shift).
  • I use the media, zoom, and window management keys all the time. I never used them on a regular keyboard because I could not remember where every keyboards keep them and they were odd key combinations (odd to me). So these mean I am using the mouse a lot less.
  • After one month, using a regular keyboard feels uncomfortable because it felt cramped and flat (I have a variety of tenting solutions) I don't remember it being much more comfortable when I started using split keyboards, but that is probably because of dealing with all the changes in geometry.
  • I noticed that I only use the sixth column on the base layer with Tab-Escape-Shift and Delete-<'>-Shift.  So I am only six combos away from switching.to a 5X3 Corne (must resist . . .)



Friday, October 03, 2025

Setting up a keymap for a Corne split keyboard to be used for data analytics

Continuing my dive into the rabbit hole known as split keyboards, I got a Corne keyboard from YMDK on Amazon.com.  A Corne is an open source keyboard (circuit board and source code are freely available. The original creator is not in the keyboard building and selling business so he lets others improve his design and sell them) It is a 3 row X 6 column + 3 (thumb row) key per side board column staggered board. The goal here is comfort from being able to independently place the halves of the keyboard under my fingers and reduce the need to twist my wrists. (i.e. reduce repetitive strain injury)

The trick with this keyboard is what to do with they symbol and control modifiers. Also known as the keymap. The answer is creating layers, such as the shift layer used for capital letters and the symbols under the number keys. So I have a symbol/navigation layer and a number/mouse control layer.

There are a few general principles I had in making the keyboard.

  1. The baseline is the standard QWERTY keyboard. I spent a lifetime building up my muscle memory and I'm not going to throw it away.
  2. I wanted to try a numpad on the right side that should also be associated with common math symbols.
  3. I wanted VI type navigation keys  (i.e. h, j, k, l correspond to left, down, up, right)
  4. Keep the symbols associated with shifted number keys on the top row, in order
  5. For symbols that did not make the base layer or the math layer, keep the symbol and the shifted symbol together (shifted symbol below the main one)y
  6. As I use it and make mistakes, move characters to the key my fingers wanted them to be.
So, the base layer is as much of the QWERTY layout as could fit.  The left column had tab above escape above shift. I started out with the escape in the corner and tab under it, but clearly my pinky wanted the tab to be next to Q. For the right column, I started out with the backspace next to P and Delete under my thumb, but a bit of use led me to switch them.  The left thumb keys had insert, enter, and lower layer (symbols and navigation), with the insert key doubling as Control when held down (also known as mod-tap). The Enter key doubled as Alt when held. The right thumb keys were raise layer (numpad, math symbols, and mouse control), space (Control when held) and backspace (alt when held). I also set up home row mods, where both hands the home row keys doubled as shift, control, alt, Command when held.  


The lower layer was for symbols and navigation.  The top row had the symbols that would normally be on the shifted number keys.  The left side had the keys that were displaced from the base layer: []{}\|`~.  The right side had the VI arrow keys on h, j, k, l. The row below those where navigation within the row: beginning of row, previous word, next word, end of row. The column to the right had page up and page down. The last column on the right had top and bottom of document (or cell for Jupyter notebooks)



The raise layer was mouse controls, numpad, and math symbols. Left side had mouse controls (left, right, up, down, clicks, scroll) and math logic symbols &|!<>.  Right side was a numpad, with -+/*_= around the numbers.  



The third or adjust layer were to control the keyboard or computer.  Left side was the keyboard, specifically lighting.  Right side were media controls, screen brightness, and screen zoom.



What really made this work was the use of combos.  I made combos (two key combinations) for the bracket type symbols which were mirrored on the left and right sides. This covered <>, (), [], {}.  In the inside columns, I made combos for :, <-, |>, # which are used in R.

Actually, this was a major modification. My initial layout had the numbers along the top row on the raise layer. But while I don't use numbers that much, when I need them I need many of them, and a numpad is less awkward.  I think.  I will know after much more use.

And here are my keyboards.  I have two Corne's, one MX with Akko Dracula (low weight linear) and XDA profile keycaps and a low profile with Kaith low profile silent tactile with low profile MX keycaps.  And a Sofle with Oetemu silent tactile switches and XDA keycaps. The Keychron K12 with Cherry Reds and OEM profile keys is what I was using before as a reference.  The Sofle is used with my work computer (having a number row is useful for making passwords smoother). The low profile Corne is packed in a bandana and a bag for travel. The MX Corne is used for my personal/non-work laptop. Since I got my first Corne in late August, this represents a big dive into the rabbit hole of split keyboards. 

For tenting, the Sofle has M5 bolts that came with it from Ergomech, the low profile Corne is using Steepy laptop risers which is what I take when out and about.  And the MX Corne is on Cooper Cases MagSafe stands. 






Tuesday, September 23, 2025

Book review: AI Snake Oil by Arvind Narayanan and Sayash Kapoor

AI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the DifferenceAI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the Difference by Arvind Narayanan
My rating: 4 of 5 stars

I read AI Snake Oil as part of the INFORMS Book Club. I work with predictive AI and generative AI at work, and I describe what I do as figuring out how AI fails, then work with my business partners to develop a process and application to make AI useful and productive. This book falls into the category of demonstrating how AI fails.

There are several chapters, each with a discussion of a way that AI fails and how the authors figured it out. But they have a pattern. First, the failures in AI is in part due to how a particular model is trained. If the training data does not match the intended use, such as the data actually represents one characteristic but the model is being used for something else. Next, they discuss that the people who made the model do not always have incentive to get it right. In particular, the large AI companies do not have incentive to either evaluate the quality of the models or improve them.

Some things I think they do well.
1. Differentiate between various generations of AI. They specifically break out predictive AI, generative AI, and symbolic AI. Each of which work differently than the others.
2. Focus on the training data. This is where AI models need to be examined (by definition, AI does not include a description of the system, so predictive and generative AI have to learn about the world through large amounts of diverse data.) And failures come from the data not matching the setting where a model is applied.
3. Be skeptical of claims that come from computer companies. I always say don't let people selling you things define terms. They also say don't let industry set the rules, the standards, or barriers of entry. Because their goal is to defend their market share, not the benefit of society.

This is a good book to read, especially as part of a discussion. Highly recommended


View all my reviews

Saturday, September 20, 2025

Business problem framing: The value of frameworks in analysis and communiction

Badge signifying completion of the INFORMS Business Problem Framing course

As part of ongoing professional education, I took the INFORMS Business Problem Framing class, which is also the lead in to the Certified Analytics Professional training that is being developed.  You may wonder what can be learned in a 5 hour course.  It is a tour of frameworks for looking at business problems.  And this is something that gets short changed in engineering training that focuses on analytic methods (i.e.math) to the expense of understanding the problem in context and working with people.  This course is needed by anyone who learned analytics as a math or computational practice and needs to learn to work with business partners/clients and is good for those who work with business partners/clients but want more range and ways of communicating (which is everyone who works in analytics in the real world)

Back when I was an enginneering professor, the supply chain center at the associated business school asked if I would be willing to coach their supply chain case study teams.  As a school with a certain amount of pretension, they found it disheartening that they were never winning these regional competitions (operations management supply chain professors and industrial engineering supply chain professors are drawn from the same pool, so this actually made sense as they needed someone who was not actually part of running a case competition.) But, after some conversations I did not approach it as a refresher on supply chain modeling, I approached our coaching sessions as teaching them how to read a case through the lens of frameworks that they had learned at some point, but may not have realized were a key tool.  Because frameworks are tools for organizing how you think of problems and a way to communicate about problems.

In every domain of expertise, experts use frameworks to organize how they look at the world.  Frameworks help experts look at all aspects of the situation.  And when provided information they can realize that they have not been given critical information that can change how they should approach a problem. Novices will often work only with what they have, like a classroom problem.  I have noted that Generative AI foundation models make the same mistake, they work with only what they have in cases where a framework would have told them a direction to go to get missing information for a fuller picture of the problem.

The other use of frameworks is communication. Frameworks provide a way of quickly communicating the essense of the issue and the information that a decision will depend on. When I was deployed in Afghanistan, I had a brief that was working its way up the chain.It was being scheduled for a 2-star general. A couple of that general's staff were present when I had given the brief at a lower level. And they told me that the brief was good, but I had to completely redo it to fit into a specific framework. Because that is how that general processed information.  (In the end, the staff went through my brief and realized that the recommendations were sound and we got the result without having to formally present the brief)

What can a 5-hour class do?  The Business Problem Framing course is a tour of a wide range of frameworks, probably familiar to someone coming out of business school, but not familiar to those whose focus is on analytical methods, math, and programming with data. And many of the frameworks are very similar in purpose.  But the right framework for the problem is on partially about the particulars of the problem, it is also about the framework that communicates the problem to everyone involved.  And just as having many ways of delivering a message is helpful to have, familiarity with a number of frameworks is helpful to have when communicating with stakeholders, because it is more possible that you will find a framework that resonates with everyone. And that leads to a better and more fruitful effort in solving the right problem at the right time.

Sunday, September 14, 2025

First days with a Corne 46 key split keyboard

I've been feeling the first twinges of soreness in my wrists, so I've started a range of actions to prevent this from getting worse.  This includes virtual physical therapy (Sword by Thryve), an ergonomic chair (off Amazon.com), a small standing desk (also from Amazon.com). And way down the rabbit hole, a Corne split keyboard. Which is notable for (1) having only 46 keys, being split, and being column staggered instead of the more traditional row staggered. (sometimes called ortholinear, but that should be used to refer to keyboards where keys are layed out in a grid, not staggered in any way.)


I bought the board and case from YMDK on Amazon.com. As my first split keyboard, I got the most basic version, a Corne  v4.1 3 X 6 + 5 key configuration (46 keys total, no screens, encoders, or other options), wired (as opposed bluetooth or 2.4Hz USB dongle) and standard MX switches. In addition I got Oetemu silent tactile switches with low profile keycaps (in purple, Go Northwestern!).

Purple corne keyboard Go Northwestern!


Actually, the first Corne I got I ordered as refurbished (i.e. used). And the right side did not work.  There is a well known failure mode and I suspect the person who had this before caused it to fail and returned it, and the Amazon.com people did not know how to test it properly when processing the return. But after some troubleshooting with YMDK I returned it and bought a new one, which worked just fine.

After making sure this worked (there is a general instruction to connect the two parts before plugging in the USB-C cable to connect it to the computer.)  Next I added switches, and finally the keycaps.  For the keycaps, the way that split keyboards work is many keys are actually accessed through layers (e.g. capital letters through a shift layer, numbers on a layer, symbols on another layer, etc.)  I touch type, so I was not concerned about knowing where the letters or default '0' layer are. So I put the keys for the left side numeber layer on the left side, and the right side symbol layer on the right side. Although, this is not much of a crutch as my keycaps are not easy to read if I don't have the RGB lighting on.

Purple corne keyboard Go Northwestern!

So the main layer is mostly the basic letter part of the keyboard, except the top row of numbers and symbols is gone. On the left side, I keep the escape key in the top left corner, and lose the Caps lock key. I get the Caps Lock key through a key combo, tapping both shift keys simultaneously.  On the right side, I have h-p then the Backspace key, so I have to move the -/_ and =/+ keys to the symbol layer.  Similarly the bracket keys and the backslash key from the next row.  The bottom row happens to fit just fine (most keyboards have a wide shift key to take up the space.)  The remaining 10 keys on each side are <ctrl>,  <layer 1>, <Enter>, <space>, <layer 2>, <Alt>  and the four center keys I mapped to <GUI>, Insert, <Menu>, Delete.  I also set up homerow mods, where keys would have one function when tapped, but another function when held.  So, the four home row keys on each side were (from inside to outside) Layer, Shift, Ctrl, Alt, Menu/GUI, with the layer and menu/GUI keys on the opposite side from where there dedicated key was found.  Not sure if I'll use these home row key mods yet (using these dual role keys is called 'tap dancing')  

What makes these small keyboards work are the additional layers. I have three: a number layer, a symbol layer, and an Alt layers.  For the number layer, I have the left side set up as a numeric keypad on the w-e-r columns. 0-+-- are on the t column, .-*-/are on the q column.  The left most column are ~, =, and <-.  The <- is because I program in R, the = needed a place, and the ~ was displaced from the number row, and is used to mean approximately or equivalent.  On the right side, the top row are the boolean (logic) operators <>$|!, for greater than, less than, AND, OR, and NOT.  For the second row, HJKL are navigation arrows for left, up, down, right.  These are the keys used by VIM (which I have set up on all of my programming environments).  Below these in the NM<> keys are beginning of line, left one word, right one word, end of line.  The two keys to the right are page up and page down. To the right of those on the edge are top of document and end of document. 

The symbols layer top row are the symbols that are shifted number keys. In the right side, the second row are the keys that were displaced: ` - = [ ] \.   The row below them are the shifted values of those same keys: ~ _ + { ] \. These were ordered so that [] and {} were in the same columns as ().  The left side second and third row has mouse navigation keys.  I actually am not sure how these work, we'll see if I figure out how to use them. :-)

The Alt layer is accessed by pressing the Layer 1 and Layer 2 keys simultaneously. The top row are the Function keys F1 through F10, with F11 and F12 continuing around the corner on the right column. The left side are keys for managing the keyboard. Specifically the keyboard lighting.  This is a wired keyboard, so there is no bluetooth to worry about. The top left button is a keyboard reset key.  The right side are audio and screen controls. Mute then volumn down and up, then media back, play/pause, media forward.  Then screen brightness down and up.  Finally Print Screen, Scroll lock, and Pause/Break.

The reason for the complications of layers is to reduce finger/hand movement. Specifically, there is no rhyme or reason for the locations of numbers and symbols, which is why people who type a lot of numbers will use a numeric keypad. So I put the numbers and basic arithmetic symbols on a number layer.  I don't have a good way of locating symbols, so I basically place them based on their normal location, with the exception of making the sets of brackets in the same columns.

For typing on the split column staggared keyboard, the basic idea is to put your hands on the table, then arrange the keyboard to be under each hand, with the columns aligned to your fingers.  However, when I do this my right hand is happy going up and down the columns. But my left hand does not want to, because it is used to having to go to the side when it goes off the home row, so I am still adjusting the left side.  The c and v keys seem to be especially problematic (my middle finger drifts right and hits 'v' when I wanted a 'c')

Just started down the rabbit hole of split keyboards., so I'll see how I master the layers. One obvious benefit, I now actually use the media keys on the keyboard, I used to always use the mouse, because I could never remember where the media key were and I did not feel like hunting for them on my keyboard.  Know, the keys are in a logical place (for my definition of logical).  I still have to think about where symbols are, but I do that for symbols that are not with the letter keys anyway. Navigation keys are easy, because I already use VIM and I did not like using the arrow keys in the lower right corner anyway.  Other potential subjects for obsession are the home row mods and key combinations. In addition to the two shift keys for Caps Lock and the F1 and F2 keys for F3, there are also common combinations for *(=[+{ that make it so these common keys can be typed without using the Layer key,

I have a Sofle on the way, and another Kickstarter backed split keyboard that I'm anticipating in November-December. This keyboard is going to be my lower profile, simple keyboard (no number row, no screens, no encoders, no bluetooth/2.4 dongle, and a 3D printed case instead of acrylic sandwhich).  I'm also exploring tenting. My simple way to start is Steepy laptop riser stand, which allow for 3 levels of elevation. I'm starting at the lowest elevation, I'll try higher options later.

Purple corne keyboard Go Northwestern!

Now back to work!

Wednesday, September 03, 2025

Subject domains that lead to failure in large language models output

At the 2025 YinzOR conference I was talking with Léonard Boussioux about types of domains where large language models (LLM) have a tendency to fail, and other conversations encouraged me to write this down.

There are stories of the early days of aviation, where a test pilot would come back and learn that his plane had cracks, and were delighted because that meant that they were learning the limits of the aircraft.  In that spirit we want to look for domains where the foundations models will give poor results, so that those developing applications can look for potential failures and design applications and train users to be attentive for errors.  For this discussion, the cause of the errors are the data used to train the foundation models.  Like other deep learning based models, to uncover categories of errors, we look at the training data.

Large language models tend to fail due to inability to work with nuance and naivete.  My friend Polly Mitchell-Gunthrie describes LLMs as unable to work with context, collaboration, and conscience.  I describe problems in LLMs as failures in nuance, naivete, and novice problems.  Again, this is due to how the foundation models are trained (effectively all publicly available text), so these are social problems, and my not be solvable in this real of LLM based AI.

Novice problems are due to the characteristics of what is available on the internet.  The majority of information on the internet is aimed at beginners. (computing topics are significant exceptions to this) So there is a lot of information that rises to the equivalent of an introductory sequence in college.  So it has a body of knowledge.  But using a body of knowledge that is targeted at introductory level leads to nuance, naivete, and novice errors.

Nuance issues are probably the most recognized.  Nuance comes into play in subjects were details matter where the answer in a specific situation is not the same as the standard case.  When given a setting, an LLM (like a novice) will take the information provided in the prompt and find other sources that include the same information and come up with an output (answer).  However, and expert would take information and fit it into an applicable framework.  Then, the expert will recognize that there is missing information that influences the final answer and ask for that information. Similarly, when considering other references, the same framework tells the expert the extent of applicability of that reference.  An LLM only matches text in the prompt with the references, so will not always check that the context of the reference matches the context of the setting of the user.  These types of issues lead experts to reach very different conclusions than people who are new to a domain, and the LLM  tend to act like novices here.  As an exercise to help people identify domains where LLMs do badly, I ask people to pick a topic that they know well, but not through textbooks or classwork, and not computer related (this tends to lead to topics that they know experiencially or through true research).  Most people identify a hobby, my manager did this exercise with his master thesis topic.  Another variation of nuance are details that frequently occur together, but are not the same.  Since the LLM works by probablistically choosing words that occur together, it can often try to combine related topics or words that should not be.  A frequent example of this is in anatomy, where LLMs trained on medical texts will often conflate the names of two body parts and into a body part that does not actually exist.

Naivete occurs when someone is in possession of facts, but does not recognize the consequences of those facts.  For an LLM, it is easy to take a prompt, then from references that match that prompt, identify other facts/details that are typically associated with the information provided by the prompt.  But unless it finds references that explicitly spell out the consequences of a particular collection of facts, the LLM will not provide the consequence.  As an example, my then 10 year old daughter had written a story that was set in a domestic setting in the United States during the 1860s (U.S. Civil War era). So when I ran her through the exercise of a topic that was not well known, she asked the Generative AI about an aspect of domestic life, specifically methods for starting fires.  Her comment was that the generative AI gave details that as far as she could tell were all true. But, it did not provide an important consequence.  When given the same set of details, a modern day chemist would mentally translate the 19th century terms to modern day counterparts, and immediately recognize that it contains all the ingredients to cause an explosion. And in real life this is what happened so there are very few examples of this technology in museums, because they all exploded. And my daughter regarded that knowing a technology meant for use in domestic (home) life had a tendency to explode to be an important detail and the LLM not reaching that conclusion to be a failure.

Another type of novice error are exceptions and crossing domains.   Many domains will teach general frameworks and rules of thumb at the introductory level.  They are intended to help practitioners succeed and to avoid common pitfalls.  However, past the introductory level practitioners learn the reasons behand the framework and rules, either from deeper training or through experience, so experts will know the exception to the rules or when to modify rules based on the particular circumstance at hand.  This is even more important in cases where multiple domains are involved, which is common outside controlled environment such as academic or teaching environments.  In this case, the standard rules for the multiple domains can conflict.  Experts will resolve this both by establishing exceptions based on the circumstance, but also looking at the ultimate goal or intent of the activity, and break or bend rules based on which rules interfere with the goals or the mission.  But they don't completely through out the rules, experts will keep in mind the intent of the rule and ensure that the intent is addressed.  When LLMs are given both the rules of the domains as well as history of prior activity, the LLMs will often identify the fact that rules are broken, and no longer follow the rules, which leads to poor outputs that do not respect the issues that arise with these domains in practice.

LLMs are especially handicapped when there are intersecting domains.  When articles or other texts are written or published, the general rule is to have anything you write/publish be on a single topic, which makes it easier to identify the target audience and for the target audience to find your work. Topics that are within intersecting domains tend to be niche topics, and are both difficult to get published and difficult to find. An thus less likely to be included in the foundation models training data.  Another area that is not found in published texts are failures.  In many domains, expertise is developed through experiencing failures. However, these domains tend not to document or publish the failures that experts learn from because of potential of repercussions or public disapproval. And if these are not published, they will not be available for training foundation models.

The purpose of this exercise is to make Generative AI useful. And to be useful the ones who work with Generative AI models have to be able to recognize and look for so that they can screen Generative AI output for other types of errors.  For example, my now 11 year old daughter continues to identify errors in Generative AI output ranging from trivial to profound, and because she has this ability, I have no concerns about her use of Generative AI.  Same with my colleagues, once they have experienced identifying errors in AI (and this holds for machine learning models as well), they are able to identify future errors and react appropriately, and not taking the outputs of AI as automatically true.  And this leads to more productive use of AI.

Sunday, August 03, 2025

Failures and how does it impact the quality of Generative AI

 I gave a talk on Generative AI as one of PyData Pittsburgh's monthly events.  While the focus of the presentation was on demonstrating impacts of randomness on Gen AI output, during the discussion we talked alot about how we teach Gen AI a specific domain, and what makes a person an expert and can Gen AI learn those things.  There were a few things about being an expert that will take a lot of work to replicate when starting with Foundation models, but one that stuck to me was the role of failure in learning,  and how hard it will be to teach this to foundation models.

My friend Polly Mitchell-Gunthrie talks about Context, Collaboration, and Consciencience when talking of the limitations of foundation models. Context is a well known, discussed, and acknowledged issue in Gen AI that we address through variations on prompting and grounding.  Conscience is both looking at issues in ethics but also mission.  But collaboration is harder, because Generative AI does not have institutional memory.  In particular, the memory of failures.

In American culture (which is where I am), we have a pressure to be perfect, and to make no mistakes and no failures.  But in a wide range of domains where there are high standards of performance, there is a maxim that is some variation of "if you have not failed, you did not try hard enough."  But even in these communities, we rarely document these failures, this level of training is done person to person, with mentors/trainers/leaders who provide cover to try different things and tolarate some level of failure in the pursuit of excellence. But more importantly for this dicussion, this does not get published, because these communities are cognizant of how intolerant of failure the general population is. But that means that the general population does not realize that the performance and excellence was developed through experiences of failure.  And the lack of documentation means that foundation models do not learn this. (for a counter example, look at baking websites that explain causes  of failure using pictures of baking disasters)

Instead of reality, the internet is a record of successes, and not failures.  This is a known problem (it is frequently discussed in academia, with journals only publishing successes, without providing lessons learned from failures, leading to a lot of wasted effort as research groups go down dead ends that other groups had already explored.) But with foundation models, that means they are trained on the successes, and not the failures.  So everything seems easy, and the Generative AI that uses these foundation models provides answers with assurance, but the people who have to implement them run in to all of the myriad of problems that come when doing things in real life.

Could you address this through grounding?  This is a cultural issue, you would need to have a record of failures, where those who went into the unknown areas of your domain were allowed to fail without adverse consequence. Then you could potentially have the Gen AI realize that a path of action could lead to an unresolved problem.  And you would have to accept the Gen AI discovering those failures, and actually telling you about them (things like this are part of the problem Gen AI has with understanding context).  So, similar to problems where there are multiple correct answers, this is as much a cultural problem in what we as a society see fit to write down (which becomes part of Foundation models), and what we do not.