Beyond Prompt Engineering: HAI, Human-in-the-Loop & Context
What we now call ‘AI art’ is derived from, or highly dependent upon neural networks, which are specific types of AI models that use interconnected nodes to process information. AI-based works, often derived from neural networks, are on the market as digital paintings, songs, books, and so on. So what happens when the output from these models or tools miss the mark? Previously, I suggested the use of human artistic intelligence or HAI, which refers to the unique cognitive and creative abilities that humans possess in artistic endeavors. In other words, we (humans) often need to use our knowledge or expertise to intervene in the AI art making process. With this in mind, I sought and found research about how human expertise could improve output from AI models.
As to our best knowledge, all of recent AI paintings involve careful curation by the technologist-artists themselves, even though the importance of this curative process is often minimized or omitted altogether. — (Chung 2021)
My daily practice of generating (and sharing) AI artwork includes making clusters of images based on the same prompts. Last week, I started working on AI-generated portraits of my favorite rap artists over the years. In some cases, very little of my intervention was needed (ex. Roxanne Shante) but I was dissatisfied with others, especially the one of Malice from the rap duo Clipse. Eventually, I cropped a photo of the rapper, edited it in Adobe Photoshop and uploaded it to the AI image generator. After a few iterations I got a satisfactory response, then I proceeded to complete the project by compositing multiple AI-generated images:
While AI can generate art based on algorithms, HAI involves a deeper, more nuanced understanding of meaning, emotion, and personal experience, making it more than just the generation of novel outputs. — World Economic Forum
The human-AI collaboration process involves humans and machines making decisions, starting with prompts provided by users, responses from machines, and further intervention, revision or fine-tuning, if needed, by both humans and machines. This process has been referred to as human-in-the-Loop or HITL, a term Neo Christopher Chung (2021) uses to explore and develop machine creativity with AI. Human creativity can not easily be codified or quantified. However, there seems to be some interest in doing just that (computationally) by marking selected AI outputs and information that can be used as attributes in further training or priming models.
Chung (and others) are exploring how to fine-tune and prime AI models:
A more complex system can be built by incorporating an additional discriminator, which is essentially a connected neural network asking if a given input has been chosen by the human curator. Interestingly, human curators may not be aware or able to explain why they have chosen certain outputs, such that the model may learn unconscious biases and emotions related to those curated set. We foresee that this will lead to new associations linking our biases, emotions, and imaginations, to visual cues, musical notes, and texts
But how would the AI model know if the output looks like its supposed to look? Or when humans may need to intervene?
LLMs, or Large Language Models, rely heavily on prompts to generate desired outputs. One of the exciting aspects of using LLMs is that the ‘human in the loop’ is simplified through text prompts, with sophisticated, multilingual language capabilities enabling artists to convey complex emotions and narratives. This is important because generative AI does produce mistakes, known as hallucinations. AI hallucinations refer to when LLMs generate incorrect, nonsensical, or misleading information while appearing to be factual and coherent. Human oversight is thus essential to correct this through reinforcement learning with feedback. For example the token prompt “Malice” did not initially generate images that were satisfactory. Through feedback, I helped the tool improve its output.
Reinforcement learning from human feedback or RLHF refines model behavior through iterative feedback loops and reward systems, teaching models to produce outputs that align with human values and expectations. Another way to guide the process is by providing the AI models/tools with contexts. Context refers to the background information and surrounding circumstances that help an AI understand a situation or task, enabling it to provide more relevant and accurate responses. Context involves giving the AI model/tool the “who, what, where, when, and why” to go beyond surface-level understanding. For example, I can use key themes or events from my favorite novels. I also composite different AI and non-AI images using Adobe Photoshop. For example, I was inspired by a scene from my favorite novel that is “Song of Solomon” by Toni Morrison.
Robert Smith, an African-American insurance agent, jumps off a roof while trying to fly as a crowd of people gather to watch. Startled by the event, two sisters drop their baskets containing velvet rose petals. Rose petals are a recurring symbol in the context of the novel’s first chapter. They are associated with artificiality, oppression, and the stifling aspects of the upper class, especially for women like the sisters.
Another image using the same method captures the dynamic between the main character Milkman Dead (brother to the rose bearing sisters) and his friend Guitar Bains. Guitar becomes consumed by his need for revenge and his belief that Milkman has cheated him out of gold. I composited two AI-generated images using a checkerboard pattern, with alternating black and white squares that commonly symbolizes duality and the interplay of opposing forces. AI can take me part of the way but my choices are made based on what I imagine and based on my experiences as a Black woman.
AI art is not simply an output of an algorithm in its raw form, but its totality which conveys — or rather create in minds of the audience — meaning and emotion. Human curators will often work with materials (e.g., what to print on) and environments (e.g., interior design), in which they put the selected AI output. Such artistic and social contexts may become valuable source materials for creative AI models. — Neo Christopher Chung
What is missing from this HITL approach is ‘human artistic intelligence’ and, as I previously stated, HAI is not easily codified and quantified. AI hallucinations and bias can ruin an art series or process. Diversity and representation is a key issue when using AI tools. Diverse artists need to be able to define high-level concepts, themes, or goals via prompts before refining specific outputs. This approach can align AI-generated output with artistic intent, offering a degree of control over outcome-based systems.