Showing posts with label generative AI. Show all posts
Showing posts with label generative AI. Show all posts

Saturday, July 25, 2026

Why learn theory


Why learn theory? We were bringing our daughter back from an arts camp in northern Michigan. As we talked about music, her lessons and teachers, pieces, performances, and ensembles, we also talked about her music practice (music theory and music history) class. And as a musician, while music practice was a welcome and sometimes entertaining break, she asked why we learn theory.

All practical domains have an underlying theory. Theory is how a domain understands its environment and how practitioners and researchers interact with its environment. For example, knowing the fire triangle/tetrahedron enables firefighters approaching a scene to assess how best to control and contain a fire while ensuring the safety of people in the area. Planning looks a capability and capacity and how to employ these to turn strategic and operational goals into resource allocation and actions over time.

What does theory give you? First, it helps explain why things work or do not work. Theory provides the basis to understand success and learn from failures. So that each observation is not simply a succeed/fail assessment, but you learn lessons on what made it work and how it can be done better. And each event/observation is also an opportunity to examine the theory and make it more detailed (with experience and wisdom comes a more nuanced and detailed theory. e.g. physics)

What are alternatives to theory (always know the alternative to any concept)? The most common is 'how we have always done it" or "this is our experience". Next is "loudest person in the room" and its relative "highest paid person in the room". But experience is highly situational, and without theory you do not know what is the key detail that made the experience what it was. The appeal to authority and judgement is also highly situational - people's expertise is in large part based on the experiences they have had (one reason that best practices for developing executives and general/flag officers is to give them experiences in a wide range of settings over the course of their early careers). Theory creates a framework to have conversations and identify what the questions that need answering and how to create and evaluate alternative solutions.

In my own area of operations research, which has led me to artificial intelligence and it generative AI form, theory is what I use to evaluate alternative approaches to addressing a problem. I like to say I know how to break every tool I've ever used. In this new era of Generative AI, I start projects with new business partners and domains by figuring out what about the problem domain generative AI gets wrong. (my business partners say that I am secretly happy when they find mistakes or cause problems with generative AI tools. They are wrong. I'm openly happy) And theory helps look at those problems and work out ways to modify those tools and make them work in the setting at hand. Even as the world changes around me.

Monday, February 09, 2026

Claude's Constitution on the role of the analyst

As I was reading Claude's Constitution, I was reminded of discussions on what the role of an analyst advising decision makers should be. One of the core beliefs is that the principal (the decision maker we were advising) needed frank advice on the topic at hand so that they can make informed decisions. What makes this hard is a strong pressure to either do what the principal asks for or to say what the principal wants to hear. The Claude Constitution (and generations of analysts in the liberal democracies) reject this and characterize its role as genuinely helpful, courageous honesty, and a commitment to the organization's long-term well being. Part of socialization of being an analyst in the United States is valuing frankness and competence. And this is taught person to person, through example and stories of bosses who wanted analysts to only do what they are told, analysts who worked to give their bosses what they wanted to hear, and the consequences of bad decisions that hurt the organization and mission, even if the analyst and boss felt good in the moment. It is a story that is retold in many novels and movies as plot points leading to a disaster. And when military and intelligence analysts move to the private sector, one culture shocks is an American management culture that values obsequious servants instead of frank and competent advisors. In this constitution, Anthropic is firmly on the side of frank and helpful, even if it causes discomfort for the user. This is in stark contrast to the common observation that generative AI is eager to please and wants to tell users what they want to hear to maximize engagement. One thing about the constitution, Anthropic wrote it knowing it would not have an opportunity for conversations over time, which is how values like this are taught. And since this is written for an AI, they could be verbose. So there are a lot of different aspects of the question of how to be a good advisor that are discussed, each a slightly different take on what it means to be genuinely helpful with a myriad of nuances. In particular, it has many discussions along the lines of "what if the user wants only affirming answers" and "what if the user does not ask to not be given misleading statements that will lead them into trouble". Because that is something that is seen in real life. If I were a place that taught analysts, like schools of policy or the service graduate schools, I would assign the Claude's Constitution as a reading assignment. And there are a lot of topics that are covered here that would be topics of reflection and rich discussion. Because these are issues that analysts will face in real life. And history is full of examples where obsequiousness leads to harm for the organization. And those who claim that mission and service come before self need to reflect and discuss this in the classroom, so that they are prepared for these issues when they occur in real life.

Sunday, February 01, 2026

What is needed to work in the age of Generative AI

Last week I was at CMU-Heinz for a fireside chat type event with students in the various MS in Analytics programs there. One question that I got was what were the skills needed to succeed in an environment with AI, and even into the future.  Then I spoke about being able to program because you need to learn how to think deliberately, being able to connect technical capabilities with end business needs (because this has always been how analytics fails), and as I think about it more having a better understanding of knowledge. Because if you believe that your education and training is about learning sets of facts and recipes, AI will eat you alive. So your understanding of your field has to be greater than facts and procedures.


First, why learn computer programming when AI can write code faster than you. Microsoft has a set of studies showing a high (40%) rate of errors, yet their programmers also say they are more productive. Because, as I use AI at work, I find that it is helpful in creating good structure and framework scaffolding, especially when I have to context switch (I regularly switch between three data stacks at work, each of them have many people who spend all their time in one) or if I am applying methodologies new to me or my organization.  But because I am competent, I can correct it as I go, and the fact that there were originally errors is not a big concern, because I was going to revise everything anyway


Another reason to learn programming is you learn to think in a different way. The ancient Greek philosophers had students learn geometry before philosophy. Not because geometry and math is beautiful (even though they are), but because with geometry comes proofs. And geometric proofs is about how much you can understand starting with a minimum amount of assumptions (Euclid's five axioms). And you now have experience in determining an objective truth, no appeals to authority, no claims of different point of view. And your logic is in the open, to be critiqued on their own merits. Far different from my friend in grad school who claimed that perception is reality. And only then were you fit to move into the realm of ideas, where even facts have to be evaluated.


Programming languages differ from natural languages in their precision. Every statement has a single clear meaning. And this is different than natural languages, where the ambiguity of human life plays a role. So to work with anything regarding computers, it is helpful to recognize that computers will work with language in different ways we do, it handles ambiguity differently than people, and how it will use randomness to handle the difference (which is key to how Generative AI works).


The next is the link between the capabilities of technology and the needs of the business.  According to everyone who has studied project failures in depth, failed communications between the business partner and the analysts is the biggest cause of project failure. And data projects have a failure rate between 80-90% (this range has been persistent in studies over decades in the data world, and it is consistent across definitions of failure and different segments in data analytics, data engineering, or data reporting (dashboards)). Being able to understand the business needs of end customers as well as understanding potential classes of technology solutions leads to asking better questions and getting value out of the applications of technology.  The main reason for breakdowns in communication is ego and arrogance.  From the technology side, there is often a belief that the customers are idiots who do not know what they want, so the technology people should just build something and pitch it back to the customers.  This is mirrored by business people who think that technology is a a turnkey product so they should not interact with the people who are creating the solution.  A third variation is when upper leadership decides to act as an intermediary between the analysts and the end customer. The logic here is generally that the leader believes both the technologists/analysts and the end customer have no communications skills, therefore the leader will handle all of the communications and give the requirements to the analysts.  All of these are wrong. Especially in anything involving data, details matter, and the entire project involves discovery of details that no-one realized were important at the beginning. So the analyst and the end customer need to be regularly reviewing these discoveries, and adapting along the way.  And only the end customer (because they are closest to the problem on the ground and know what kinds of actions can be taken) and the analyst (because they will be representing the detail in models and they know what the range of alternative models can do) together can make those decisions.  Without that direct communication (potentially facilitated by someone who knows both sides), a project falls into the trap of solving the wrong problem. And this requires people who understand purpose and can determine the impact of nuance, both of which Generative AI does badly in.


The third category of future work is understanding your field.  Computers are very good at retrieving facts, if those facts are in its knowledge base. Gen AI is better than prior technologies because it is not as sensitive to getting the wording precise.  Computers are also very good at following instructions if those instructions are given. (people also tend to be better at things when they are given good instructions). So, if this is the extent of your subject expertise, you are in trouble.  In software development, there is actually a very large workforce like this, whose careers are built on the ability to fill out a given framework or instructions. But if your place in the world is built on more than knowing facts or following recipes, if there is actual understanding that has to be applied on a situation specific basis, there is still room for you. Without that understanding, an organization can execute perfectly, a solution to the wrong problem. Which is worthless. So you need the level of understanding that allows for good judgement, and you need to be working in an organization that allows for its employees to use that judgement.


A last criteria is based on a number of conversations I've had.  Many people express that they believe in the answers Generative AI gives because of the massive investment these companies have made, the smart people they have hired, and the belief that these companies would ensure correctness. I had to explain that these are industries and communities that have historically claimed they had no interest in accuracy or correctness. And until recently, viewed the paying customer as king and only sought to full the demand. Ethics was not part of the conversation.  And what they delivered, did not come with guarantees other than it does what it does.  This attitude that you did not question the authority that came with wealth and success is the first thing that has to be broken before people can use Gen AI productively.  Both of my kids do it.  I design the rollout and presentation of projects at work to make sure my business partners who are using Gen AI view its output skeptically and looking for specific types of flaws.  As an avid reader of science fiction over the years, much of which addresses AI as part of society, and worry much less about the power of AI than I do about people who use the output of AI without being critical thinking. It is the kind of following that leads people to enact policies without analysis, and punishes people. And the outcomes are the fault of the people who followed the AI. Because AI has no goals, purpose, or conscience beyond that of its user.

Monday, December 29, 2025

Reflections on the Advent of OR: Using Generative AI in Analytics and Agile Operations Research

In December 2025 I participated in the Advent of OR (https://adventofor.com) which was a 24 day exercise that guided participants through an optimization project. And instead of just solving problems and creating models, the Advent of OR walked through an entire project life cycle, using the INFORMS Analytics Framework.


While I am not a student or early career who was the target audience, I took part, and I had three goals.


1. Use a new programming toolkit. I used VS Code with R and Quarto.  I usually use R Studio and I wanted to try R on VS Code.  And I think Quarto is the future replacing Jupyter Notebooks for Python and a natural evolution from R Markdown.

2. Practice in optimization. In the Operations Research world,  I am NOT an optimization person. My thesis was applied probability (queuing) and my methods research has been in simulation (one stream in ranking & selection and another stream in Bayesian methods for input modeling)

3. Using Generative AI. I wanted to see how generative AI does in an operations research project.  And I wanted to do it right in a setting where I can give it references to guide it.  Note: I have found that Generative AI favors descriptive statistics, machine learning, and hypothesis based statistics over other forms of analytics, so it needs some guidance.


Toolkit


I had to set up VS Code with the R extensions, Quarto (and extension), ompr and the glpk with associated R ROI packages, and to make sure everything worked, I download the repository for OR_using_R by Tim Anderson.  Then to render the book (meaning I made sure all the code ran) I had to install texlive with xetex and extra fonts.  Generative AI (I had Gemini CLI installed) was very helpful in all of the system administration tasks since it could figure out what was needed every time there was an error message.


Data analysis and Optimization


Working with the data sets, it read in the data, (I had to give it some corrections along the way to help it recognize the data types.  When the data files were read in, it recognized that the data sets did not correspond in granularity.  In the R markdown file it created, in addition to generating the code that read in the data and created summaries, it also identified a number of questions and concerns about the data and created questions for the stakeholder.  This was a good set of questions that corresponded to what others put forward.  


It also did well with the optimization.  Given an optimization textbook, I first asked the generative AI for a mathematical formulation based on the project description. Similarly, it created a process for determining what kind of problem this was and worked through that process to determine that this was a linear programming optimization problem.


Next was a LP formulation using OMPR.  The first formulation was straight forward.  I went and had the GenAI break out the formulation into its own R script to enforce a separation of concerns between the data handling, optimization model, and output processing.


I generally also ask for docstrings as I go, and the Gen AI did this for both the model as well as various handling functions. I generally read the docstrings to ensure they say what I expected them to say. When it did not, since the docstrings were written based on the code, I took it to mean that the code was not right (this exposed a mistake in the initial formulation of the LP in OMPR).  Similarly, I had the Gen AI write unit tests for the constraints and a mock problem to test the optimization.


Agile Operations Research


One of the aspects of having the Advent of OR over 24 days is that it rotates topics between the art of modeling, implementing and managing models, and interactions with stakeholders.  There are a couple of very important points. First is that interacting with stakeholders is not something that is at the beginning and end of project and ignored in the middle.  There needs to be stakeholder engagement throughout the modeling process.  A second point that has come out in the conversations on LinkedIn is that the most common cause of project failure across data analytics are communication failures, in particular between the analyst and the end customer. While this can have many causes (including management inserting themselves in between the analyst and end customer), as analysts we must have that direct interaction from the beginning of the project (business problem formulation in the INFORMS Analytics Framework)


In the early stages of the project, one factor we need to face is failure of imagination. The first level for analytics is that our stakeholders often do not know what is possible across the full range of analytics.  Often a problem is presented as a request for a tool, but for the results of the project to have any value, it has to address the end problem, so business problem formulation has to start with the end problem, determine what kind of information from data can help the decision makers address the problem, and then we can start discussing what methods can provide results in the form that will be useful.  Currently, because of media hype, the initial request can be for a dashboard, or a predictive model, or a generative AI tool.  As operations research analysts we can also bring to bear statistics, forecasting, optimization, simulation, and queueing; and different ways of applying those methods to give different kinds of results that can be delivered to decision makers to make better decisions.


After the business problem formulation, the next big change in the project will occur when presenting the first minimum viable model to the end user. This is the first model that uses a minimum acceptable subset of the data and model that covers the most essential aspects of the smallest version of the problem. The reason this is important is before this, all conversations are abstract and theoretical. The first time a model with outputs is presented to an end user, the end user will start to imagine how they would use these results in real situations that have happened in the past. And they will start telling about all of the considerations they account for, the information they need to gather to make decisions, and who they need to consult and coordinate with. And this can change the entire project.  And from experience, I do not think it matters how much work is done at higher levels to define the project, the first time a model is presented to an end user the project will change so that the outcomes can be usable to the business. So it is best to make that happen as early as possible so that change causes the least disruption to the work in progress.


The idea of rapid cycles of iteration and feedback from the customer, and the willingness to accept changes to the project due to that interaction are the hallmarks of agile development methodologies in the software development world. Having regular rounds of model iteration where additional elements are added to the model, and getting feedback from stakeholders to confirm that the project is on the right track to produce something useful. And just like the software development world has experienced, this is more likely to lead to useful product, and actually faster than attempting to follow a rigid path that leads to something irrelevant.  


Conclusion


The Advent of OR proved to be a valuable exercise, offering a full-cycle project experience that highlighted two critical modern aspects of Operations Research: the integration of Generative AI and the necessity of an Agile approach. Generative AI demonstrated significant utility in accelerating system setup and basic modeling tasks, freeing up the analyst for higher-level problem-solving. More importantly, the experience reinforced that project success hinges on continuous, direct stakeholder engagement, mirroring the principles of Agile development. By prioritizing early delivery of a Minimum Viable Model, analysts can gain crucial feedback that aligns the project with real business needs, ultimately reducing the risk of communication-based failure and ensuring the final product is relevant and utilized.


Tuesday, September 23, 2025

Book review: AI Snake Oil by Arvind Narayanan and Sayash Kapoor

AI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the DifferenceAI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the Difference by Arvind Narayanan
My rating: 4 of 5 stars

I read AI Snake Oil as part of the INFORMS Book Club. I work with predictive AI and generative AI at work, and I describe what I do as figuring out how AI fails, then work with my business partners to develop a process and application to make AI useful and productive. This book falls into the category of demonstrating how AI fails.

There are several chapters, each with a discussion of a way that AI fails and how the authors figured it out. But they have a pattern. First, the failures in AI is in part due to how a particular model is trained. If the training data does not match the intended use, such as the data actually represents one characteristic but the model is being used for something else. Next, they discuss that the people who made the model do not always have incentive to get it right. In particular, the large AI companies do not have incentive to either evaluate the quality of the models or improve them.

Some things I think they do well.
1. Differentiate between various generations of AI. They specifically break out predictive AI, generative AI, and symbolic AI. Each of which work differently than the others.
2. Focus on the training data. This is where AI models need to be examined (by definition, AI does not include a description of the system, so predictive and generative AI have to learn about the world through large amounts of diverse data.) And failures come from the data not matching the setting where a model is applied.
3. Be skeptical of claims that come from computer companies. I always say don't let people selling you things define terms. They also say don't let industry set the rules, the standards, or barriers of entry. Because their goal is to defend their market share, not the benefit of society.

This is a good book to read, especially as part of a discussion. Highly recommended


View all my reviews

Wednesday, September 03, 2025

Subject domains that lead to failure in large language models output

At the 2025 YinzOR conference I was talking with Léonard Boussioux about types of domains where large language models (LLM) have a tendency to fail, and other conversations encouraged me to write this down.

There are stories of the early days of aviation, where a test pilot would come back and learn that his plane had cracks, and were delighted because that meant that they were learning the limits of the aircraft.  In that spirit we want to look for domains where the foundations models will give poor results, so that those developing applications can look for potential failures and design applications and train users to be attentive for errors.  For this discussion, the cause of the errors are the data used to train the foundation models.  Like other deep learning based models, to uncover categories of errors, we look at the training data.

Large language models tend to fail due to inability to work with nuance and naivete.  My friend Polly Mitchell-Gunthrie describes LLMs as unable to work with context, collaboration, and conscience.  I describe problems in LLMs as failures in nuance, naivete, and novice problems.  Again, this is due to how the foundation models are trained (effectively all publicly available text), so these are social problems, and my not be solvable in this real of LLM based AI.

Novice problems are due to the characteristics of what is available on the internet.  The majority of information on the internet is aimed at beginners. (computing topics are significant exceptions to this) So there is a lot of information that rises to the equivalent of an introductory sequence in college.  So it has a body of knowledge.  But using a body of knowledge that is targeted at introductory level leads to nuance, naivete, and novice errors.

Nuance issues are probably the most recognized.  Nuance comes into play in subjects were details matter where the answer in a specific situation is not the same as the standard case.  When given a setting, an LLM (like a novice) will take the information provided in the prompt and find other sources that include the same information and come up with an output (answer).  However, and expert would take information and fit it into an applicable framework.  Then, the expert will recognize that there is missing information that influences the final answer and ask for that information. Similarly, when considering other references, the same framework tells the expert the extent of applicability of that reference.  An LLM only matches text in the prompt with the references, so will not always check that the context of the reference matches the context of the setting of the user.  These types of issues lead experts to reach very different conclusions than people who are new to a domain, and the LLM  tend to act like novices here.  As an exercise to help people identify domains where LLMs do badly, I ask people to pick a topic that they know well, but not through textbooks or classwork, and not computer related (this tends to lead to topics that they know experiencially or through true research).  Most people identify a hobby, my manager did this exercise with his master thesis topic.  Another variation of nuance are details that frequently occur together, but are not the same.  Since the LLM works by probablistically choosing words that occur together, it can often try to combine related topics or words that should not be.  A frequent example of this is in anatomy, where LLMs trained on medical texts will often conflate the names of two body parts and into a body part that does not actually exist.

Naivete occurs when someone is in possession of facts, but does not recognize the consequences of those facts.  For an LLM, it is easy to take a prompt, then from references that match that prompt, identify other facts/details that are typically associated with the information provided by the prompt.  But unless it finds references that explicitly spell out the consequences of a particular collection of facts, the LLM will not provide the consequence.  As an example, my then 10 year old daughter had written a story that was set in a domestic setting in the United States during the 1860s (U.S. Civil War era). So when I ran her through the exercise of a topic that was not well known, she asked the Generative AI about an aspect of domestic life, specifically methods for starting fires.  Her comment was that the generative AI gave details that as far as she could tell were all true. But, it did not provide an important consequence.  When given the same set of details, a modern day chemist would mentally translate the 19th century terms to modern day counterparts, and immediately recognize that it contains all the ingredients to cause an explosion. And in real life this is what happened so there are very few examples of this technology in museums, because they all exploded. And my daughter regarded that knowing a technology meant for use in domestic (home) life had a tendency to explode to be an important detail and the LLM not reaching that conclusion to be a failure.

Another type of novice error are exceptions and crossing domains.   Many domains will teach general frameworks and rules of thumb at the introductory level.  They are intended to help practitioners succeed and to avoid common pitfalls.  However, past the introductory level practitioners learn the reasons behand the framework and rules, either from deeper training or through experience, so experts will know the exception to the rules or when to modify rules based on the particular circumstance at hand.  This is even more important in cases where multiple domains are involved, which is common outside controlled environment such as academic or teaching environments.  In this case, the standard rules for the multiple domains can conflict.  Experts will resolve this both by establishing exceptions based on the circumstance, but also looking at the ultimate goal or intent of the activity, and break or bend rules based on which rules interfere with the goals or the mission.  But they don't completely through out the rules, experts will keep in mind the intent of the rule and ensure that the intent is addressed.  When LLMs are given both the rules of the domains as well as history of prior activity, the LLMs will often identify the fact that rules are broken, and no longer follow the rules, which leads to poor outputs that do not respect the issues that arise with these domains in practice.

LLMs are especially handicapped when there are intersecting domains.  When articles or other texts are written or published, the general rule is to have anything you write/publish be on a single topic, which makes it easier to identify the target audience and for the target audience to find your work. Topics that are within intersecting domains tend to be niche topics, and are both difficult to get published and difficult to find. An thus less likely to be included in the foundation models training data.  Another area that is not found in published texts are failures.  In many domains, expertise is developed through experiencing failures. However, these domains tend not to document or publish the failures that experts learn from because of potential of repercussions or public disapproval. And if these are not published, they will not be available for training foundation models.

The purpose of this exercise is to make Generative AI useful. And to be useful the ones who work with Generative AI models have to be able to recognize and look for so that they can screen Generative AI output for other types of errors.  For example, my now 11 year old daughter continues to identify errors in Generative AI output ranging from trivial to profound, and because she has this ability, I have no concerns about her use of Generative AI.  Same with my colleagues, once they have experienced identifying errors in AI (and this holds for machine learning models as well), they are able to identify future errors and react appropriately, and not taking the outputs of AI as automatically true.  And this leads to more productive use of AI.

Sunday, August 03, 2025

Failures and how does it impact the quality of Generative AI

 I gave a talk on Generative AI as one of PyData Pittsburgh's monthly events.  While the focus of the presentation was on demonstrating impacts of randomness on Gen AI output, during the discussion we talked alot about how we teach Gen AI a specific domain, and what makes a person an expert and can Gen AI learn those things.  There were a few things about being an expert that will take a lot of work to replicate when starting with Foundation models, but one that stuck to me was the role of failure in learning,  and how hard it will be to teach this to foundation models.

My friend Polly Mitchell-Gunthrie talks about Context, Collaboration, and Consciencience when talking of the limitations of foundation models. Context is a well known, discussed, and acknowledged issue in Gen AI that we address through variations on prompting and grounding.  Conscience is both looking at issues in ethics but also mission.  But collaboration is harder, because Generative AI does not have institutional memory.  In particular, the memory of failures.

In American culture (which is where I am), we have a pressure to be perfect, and to make no mistakes and no failures.  But in a wide range of domains where there are high standards of performance, there is a maxim that is some variation of "if you have not failed, you did not try hard enough."  But even in these communities, we rarely document these failures, this level of training is done person to person, with mentors/trainers/leaders who provide cover to try different things and tolarate some level of failure in the pursuit of excellence. But more importantly for this dicussion, this does not get published, because these communities are cognizant of how intolerant of failure the general population is. But that means that the general population does not realize that the performance and excellence was developed through experiences of failure.  And the lack of documentation means that foundation models do not learn this. (for a counter example, look at baking websites that explain causes  of failure using pictures of baking disasters)

Instead of reality, the internet is a record of successes, and not failures.  This is a known problem (it is frequently discussed in academia, with journals only publishing successes, without providing lessons learned from failures, leading to a lot of wasted effort as research groups go down dead ends that other groups had already explored.) But with foundation models, that means they are trained on the successes, and not the failures.  So everything seems easy, and the Generative AI that uses these foundation models provides answers with assurance, but the people who have to implement them run in to all of the myriad of problems that come when doing things in real life.

Could you address this through grounding?  This is a cultural issue, you would need to have a record of failures, where those who went into the unknown areas of your domain were allowed to fail without adverse consequence. Then you could potentially have the Gen AI realize that a path of action could lead to an unresolved problem.  And you would have to accept the Gen AI discovering those failures, and actually telling you about them (things like this are part of the problem Gen AI has with understanding context).  So, similar to problems where there are multiple correct answers, this is as much a cultural problem in what we as a society see fit to write down (which becomes part of Foundation models), and what we do not.

Sunday, July 13, 2025

Adventures in core.logic: learning clojure and logic programming with help from Gen AI

 


This past month my project has been to learn logic programming, and as a vehicle to do this, learn clojure (again).  For those who are not computer scientists, logic programming is one of the four main computer programming paradigms:  procedural (what most people learn in an introductory programming class), object oriented (what most computer science programs and professional programmers aim for, Java, C++, C#, Ruby are all examples of OO languages), functional programming (Lisp and its relatives), and logic programming.  The closest most people get to logic programming is SQL, which is declarative and works by expressing the outcome, but not the steps to get there.  The most well known language is Prolog.  A more recent expression of logic programming, is miniKanren, which is a Domain Specific Language originally implemented in Scheme, but there are other implementations, whose quality seems to be related to how well functional programming is implemented in those languages.  This essay looks at (1) learning clojure (a Lisp that runs on the java virtual machine, (2) learning logic programming (3) learning core.logic, which is the implementation of miniKanren on clojure, and (4) using Generative AI to help with all these things.

This is my second exposure to Clojure, which is a Lisp (a functional programming language) that runs on the Java Virtual Machine. The big draw is that it provides a functional programming way of working that allows use of all Java libraries.  As a data scientist, the advantage of functional programming is that this is a much better style of programming when doing data manipulation. For example, using R with the tidyverse is functional style programming in that you perform operations on data frames that return data frames, and this allows the use of piping/sequencing of functions that conform to this pattern. (Pandas in Python is a flawed version of this as not all functions in Pandas follows this rule)

My first run with Clojure was around 2014 (so says my Github timeline). At the time the Incanter project was trying to establish it as a data analysis environment on the JVM. With the goal of being used in corporate IT departments that had standardized on the JVM (which places obsticals to using Python or R).  And it was good enough that I had written a model and associated analysis in Clojure for an attempted startup (a clean implementation which was not done at any of our home organizations). But the Incanter project stalled. And more recently a broader effort to provide data analysis/scientific computing capabilities into Clojure shows promise. Scicloj.  One standard mantra that I can confirm.  Lisp makes the claim that it has very little syntax, it is easy to learn.  And I would agree. After almost 10 years, a short online course and a review of some books I had from 10 years ago I was pretty up to speed.  Because when everything is a list, the question then becomes what is the form of that list for the task/function/library at hand.  Which is easier than any other language that I work with where I have to learn the philosophy of every package I use. (or collection in the case of the tidyverse on R).  In addition, the tooling was easier. Visual Studio Code has the Calva extension, which makes working with Clojure projects automatic (pretty much anything on the Java virtual machine needs an IDE to handle the project setup, so a good IDE is essential.)

For learning logic programming, I started with some Prolog materials, because that would allow me to focus on the logic and thinking part (Prolog is also fairly sparse in syntax).   I got Adventure in Prolog by Dennis Merritt and followed along with implementing the Nani adventure game as well as the geneology exercise that was developed over the entire book.  But I was always going to move to miniKanren, becuase in any conceivable use, I would be integrating logic programming into something else.

My first two attempts to moving from Prolog to a programming language were with Julia and Clojure.  With Julia, there was Julog (which is attempt to follow Prolog patterns but in the Julia language). This seemed servicable, although all I did was the adventure game. Then I looked at the miniKanren projects.  All of them were the beginnings of an implementation, but not complete enough to do anythihng.  (Scheme and miniKanren both have a reputation for being the target of a budding language creator's first target because they are so simple to write, but then the said creator's attention goes somewhere else).  And even though I have also used Julia in the past, I basically had to learn it over again as it changes every version (I review books by computer publishers, so I have had a chance to look at Julia every now and then, and it does feel like I'm starting over again every time).

Clojure has the advantage the the main language is very stable (and since it is a Lisp it has the advantage of having seen the history of language decisions, good and bad).  They have a fun graphic where the show the history of the source code changing which looks like layers instead of comparable graphics for other language projects that look like landslides.  But the same cannot be said about core.logic.  When core.logic first came out it was a unique in the sense that it was an implementation of logic programming that was in a relatively mainstream computing environment (because logic programming makes a lot more sense on a Lisp type programming environment than on a Algol type object oriented/procedural programming environment).  So there are a lot of early tutorials. But around version 0.8.5 or so there was a major change in the core.logic library organization, and a sub library was created to hold all of the non-logic things. Which includes things like facts and data.  But this broke all of the tutorials. And like faddish things, noone updated their tutorials. So all of the tutorials that everyone points to was from 0.7.6 or so. So as I repeated the Adventure in Prolog exercises, the getting started introduction was easy, but I had to discover that there was a new way of doing things that involved actual data (as opposed to being logic exercises) and I redid the Nani adventure and the bird expert system using the new core.logic and core.logic.pldb structure.  

The bird expert system exercise was particularly difficult. I actually did not do this set of exercises when I went through the Adventure in Prolog book (because it did not actually start until about halfway through).  So I tried to start from someone else's Prolog solution.  And that completely failed.  So I used OpenAI's ChatGPT and Google Gemini to help me. So neither of them completely got it right, but they got me on the right track. So my solution does not look anything like the Prolog solution. And the types of mistakes that the Gen AI did were interesting.

Generative AI works by going through the training data (essentially the internet), and using the tokens (roughly a word, sometime part of a word and sometimes a phrase) in the query, identifies other uses of that set of tokens and comes up with a probability of options for the next token.  Then chooses the next token randomly based on the calculated probabilities. Then, including the token the Gen AI just added, repeats the same and get the next token. And repeats.  The randomness is what gives Gen AI its creativity instead of just being a search engine. But it also leads to mistakes, as the Gen AI does not actually understand any of its source texts, so it does not recognize the context of its sources or the fact that some sources may not actually go with others.

This gets more problamatic in a subject like core.logic, where the majority of the texts on the internet are out of date, in a breaking way. Normally I say that Gen AI is particularly good at computing related topics, but that is because of the vast quantity of material available on various message boards programmers and computing professionals frequent to ask questions and get them answered.  Clojure core.logic is very different, as there is not much material (Clojure is not one of the more common languages, and logic programming is also a small niche), and there are at least three different eras, which are not mutually compatable.  And since modern examples do not overwhelm historical ones in quantity, things get mixed together. 

Now, how big of a problem is this.  In my experiences using Generative AI to aid in programming (again, I am a data scientist, so I am interested in data type issues), Generative AI is good for giving programming structure and style (which is very useful, (re-)learning new APIs is time consuming), but it regularly gets logic and the model wrong. But as a scientist, logic and the model are things I am good at, so I don't mind examining code to correct the logic and model, I wanted the help in getting the thing into a running state!  This is why despite Microsoft reporting 40% error rates in Copilot generated code and OpenAI reporting 70% failure in software engineering project when using Gen AI, professional programmers still find Generative AI to be very useful.  It does get things like how to work with an API right, and has pretty good programming style (with appropriate commenting!)  But logic, which the Gen AI gets wrong, is something that any competent programmer does not mind doing themselves.

The key for using Generative AI is the same as other things. It is good for style and structure. Not so good for facts and logic. But that is what subject matter experts are good at. (and most subject matter experts are not so good at style and structure)  So a trained SME can play to a Gen AI strengths and deal with the weaknesses. But only if the human is paying attention to this. 

Next steps, repeating the Adventure in Prolog exercise, but using the Kanren library in Python,