26 Comments
User's avatar
Nour's avatar

Shallow intent is a great way of putting it. I think another part of it is that people don’t like to feel that they are being played or lied to. We consume content we care about; we want authentic, trusted voices and expect effort. When someone realizes a piece of content they’re consuming is AI-generated, because of these markers of shallow intent, they see it as a sort of betrayal of trust that real effort was put into it, and thus perceive it to have less value.

Matt's avatar

The only one I would buy is Fulvio Bisca’s dragon. This one actually looks like a poster not AI and is a good talking point about UT in general.

AKCH Haine's avatar

True. But what a wealth of beautiful and captivating designs that were submitted. Some true artists here.

Matt's avatar

Yeah for sure definitely good art and/or graphics. Just saying if the goal is a poster product as merch that people may put up in homes or dorm rooms or what have you then it’s the dragon.

Ron Bodkin's avatar

I really appreciate your analysis of shallow intent.

The space where people have made the most progress on deep intent with the ability to control and maintain coherence is in software engineering. This still requires active human steering, but it also has evolved a lot of craft (including an increasing emphasis on having AI interview the SME extensively to get clear requirements). I am struck by how "uncontrollable" AI image generation is (and indeed to control AI writing you have to do the same work as writing - specifying exactly what each small part needs to say, revising it etc). I agree that AI advances it may well become more coherent and able to hold and execute deep intent.

I think there's also a point about coherent style - an uncanny valley tell that people react to (as in other commenters saying that people feel lied to/want to connect with authenticate human expression). I think this will also bifurcate - for functional purposes AI output will be embraced but for art/enjoyment it will be harder to convince people.

Fukitol's avatar

You have nailed the central problem but it's also the reason LLM art will *never* be good. Art is an attempt at communication. "A picture is worth a thousand words." The non-linguistic medium communicates things that reside in the non-linguistic parts of our mind. For communication - information transfer - to take place, you need an encoding of an intent to communicate, and a decoder that understands the encoding method and can unpack the content. The more successful this process, the more powerful the art.

What can a thing with no mind communicate? Its can be interesting, aesthetically pleasing even, but ultimately of no more interest than the mandlebrot set. All the intent that was ever in it is what is written in the user's prompt. Stochastic fill has no intent; it's noise. Noise can be aesthetic - film grain, paint spatter - but never adds information, only removes it in attempt to enhance aesthetic.

So your LLM generated image is not worth a thousand words, it's worth a hundred, less the words lost to the algorithm and stochastic drift. You create something even shallower than your original thought. Your Mesopotamian image wastes a book worth of potential information to communicate a summary caption worth of thought. You wasted the viewer's time, frustrated his attempt at interpretation, asked him to dive head first into your puddle.

More elaborate noise algorithms won't fix this. You need a mind, not a crude mathematical facsimile, to make art. If there's a way to give magic sand a mind, we haven't figured it out yet and there's no guarantee we will.

Tomas Pueyo's avatar

I hear you, but for example you could have a prompt that says "create a Mesopotamian family in 2000 BC", and instead of generating the image, it basically asks ChatGPT all the elements that should go in that, and then an agent loop to fact-check it, and then another agent loop to push it aesthetically, and from these loops you get a very long and detailed prompt, which you then feed into the image generator, which then goes through loops of fact-checking, until it gives you something precise and full of intent, as in every element is truthful and very telling of how life was there back then.

That sounds to me like a way to get massive amounts of intent, even if not human.

Fukitol's avatar

You're looking for a loophole but you just described a process to generate more stochastic filler. I'll grant it might be more accurate filler (although current gen diffusion models are not up to the task, try writing a very detailed prompt sometime). But the intent is still just your original prompt, the rest is improving the fill.

And it'll always be slightly deranged and uncanny, just because much of what is art is non-linguistic and derived from human sensory experience. LLMs are language models, and so at best crudely simulate the little nugget of tissue in our left hemisphere devoted to linguistic symbol manipulation. They're not even good at simulating adjacent regions used for logic and mathematics.

So a language model that makes images can associate image features with words, and on the one hand that's kind of amazing, but on the other hand most of what is in art is not easily expressed in words, thus opaque to a language model.

Think of it this way: you could write an entire book describing in minutia the features of a famous painting. Hand that book to a good artist and ask him to reproduce the painting. You won't get a good copy. On the other hand show the artist the painting once, let him study it, then ask for a copy and he'll produce something that, if not identical in style or stroke, captures the essence of it.

Funny example, the corporate memphis version of Goya's Saturn that made the rounds a while back. Not only did it capture the essential theme and horror of the original, but added a whole new layer of commentary.

Why is that? Because there's information there that does not readily translate to words. So the model can't grok it, because it's a model of word associations. Of course it knows the features of an image called “Saturn Devouring His Son” and can produce a million variations of it, but can never do so meaningfully.

So, could we make a non-linguistic model of art? Maybe, with enough compute. You'd be simulating a much larger portion of the human brain so come with orders of magnitude more hardware at the ready. But if you could, then it could grok art. But at the same time, you'd not get to give it a one-liner prompt. Or you would, and get the same result as doing the same with an artist - something unexpected and no longer really yours.

chayote tacos's avatar

A brave post. Well executed and stated

i nompelis's avatar

Igor 4th.

Peter Scheffler's avatar

Thanks. I think you have hit on something important. AI as execution is helpful, but the experience and awareness of the creator are essential. Without that, the product is likely to be trivial eye candy and not effective or moving. So in the rush to adopt AI, companies run the risk of eliminating the jobs of the people we need to be using it. To me, it's a bit like the rise of electronic music and turntabling. People can easily create sounds and effects, but without real understanding of music, the results often have been forgettable. That's okay, it can help develop the experience and awareness of the creator with proper feedback like yours, but it shouldn't be assumed to be good at the beginning, just because it's AI.

Tim Lundeen's avatar

I've started using AlterAI instead of the better-known options. It is much less biased by mainstream "consensus", and has an option to reduce hallucinations by checking/merging search results with the training data. I'm impressed, and highly recommend it. https://alter.systems/

RenOS's avatar

AlterAI unfortunately seems to be backed in a similarly one-sided, albeit different, manner than the mainstream; All four "what people are saying" blurbs they present are from vaccine-sceptics. I'm relatively critical of the covid policies myself so I have no problem if sceptics are treated fairly, but quoting *only* them does not seem balanced, either.

I've tried it with a few different prompts on different topics I consider genuine controversies and others I'd consider debunked. To its credit, it does not seem to get the facts wrong, but it writes absolutely every story in this grating "what the mainstream isn't telling you" voice. Again, there's lots of things the mainstream isn't telling you so I have little problem in principle, but that's just a really limited perspective.

It also completely failed on a few queries that Chatgpt had no issue on.

Tim Lundeen's avatar

Good feedback, thanks

Emma's avatar

oh trust me we do mind when something is made with AI BECAUSE it’s made with AI

Janet O'Keeffe's avatar

Frankly, I don't like any of them they're all too dense and overly complicated.

Theodora's avatar

I’m immensely grateful for the opportunity to have participated in this competition, and also to have my entry spotlighted. Thank youuuuu!

Great job to everyone that participated! So many amazing entries were actually submitted and congratulations to the winner, Igor! Totally deserved it. 👏🏾 🤍

Jaime's avatar

Very interesting article! I really liked how you described AI Grime and, more broadly, your idea of intent density. I think it's a really useful way of thinking about AI-assisted creative work, and it's definitely something I'll keep in mind for future designs.

Igor's posters are absolutely fantastic. Do you know if he has a portfolio, repository, or social media account where I could follow his work? They struck me as a really effective way of drawing people into reading an article, they're visually compelling while also distilling the core idea into a single memorable visual. That's exactly the kind of work I'd love to learn from.

Omercan's avatar

Also, you stated that the entire design should not be made with artificial intelligence, and it was quite interesting that the winning work was made with artificial intelligence. I am waiting for your feedback on this subject. It is really an artist. Thank you for your understanding.

Omercan's avatar

Hello, I'm Ömercan Göktürk, I couldn't see any comment on my poster among the comments you made, and I didn't use artificial intelligence during the drawing phase of the poster, I'm waiting for your comments on this subject.

JustRaven's avatar

Tomas did end his newsletter with this remark, which you might have overlooked:

"Once again, thanks to all for your submissions! Apologies if I didn’t get a chance to talk about yours, there were too many to go through one-by-one. Rest assured that I looked at every poster carefully. Please keep creating, you all put yourselves out there by entering this, and that in itself is to be commended!"

cwiles's avatar

I like the vikings as a powerful image out if nowhere

Sascha Klocke's avatar

I wonder to what extent intent plays a role, versus just being lucky in your rolls of the dice to get something with fewer (obvious or not so obvious) errors and then more success in the re-prompting to fix them without introducing a whole bunch of new errors elsewhere.

I think the winning entries are still crowded with random elements, lots of errors, unclear design decisions, scale and perspective mismatches, curious new chess pieces, etc.