Training AI to Paint with Code

Training AI to Paint with Code

A collaborative project between the author and a colleague named Cameron has produced a reinforcement‑learning system that teaches a language model to generate images by writing editable JavaScript code instead of relying solely on textual prompts. The model receives a description such as “draw a peach hibiscus in watercolour,” composes a complete p5.brush sketch, renders it in a sandboxed Puppeteer environment to create a PNG, and then receives a reward based on how that image compares to two randomly chosen reference paintings drawn from a hand‑rated pool. The loop—prompt, code generation, rendering, judging, reward, and model update—is executed thousands of times during training, allowing the code itself to become the editable artefact that can be refined without re‑prompting the AI.

The key breakthrough emerged when the researchers re‑engineered the reward rubric. The original system used nine separate signals, many of which were highly correlated, causing the model to plateau at a 0.65 reward with repetitive, flat‑clip‑art outputs. By replacing absolute scoring with pairwise judgments—asking a judge model which of two images better matches the prompt—and consolidating the rubric into four components (a compile‑and‑brush gate, a length check, the HPSv3 human‑preference model, and the pairwise judge weighted at 60 %), the reward signal gained dynamic range and variance. This change allowed the model to break the plateau three times faster, continue improving, and dramatically reduce code length from 13,500 to under 2,000 tokens, demonstrating that concise code can still produce high‑quality compositions.

The experiment highlights both the promise and the challenges of applying reinforcement learning to creative design tasks where aesthetic quality is subjective. While the revised rubric yielded faster convergence and richer outputs, the authors note that a next step—training a dedicated reward model on the hand‑rated pool (proper RLHF)—was not yet pursued, which would enable the system to assess quality without constant pairwise comparisons. The approach opens a pathway for more controllable AI‑generated art, offering users the ability to edit generated code directly, and suggests that careful construction of reward functions is crucial for advancing AI creativity beyond static prompt‑based generation.

Sources cited: 📰 Hacker News ↗

⚡ Effects Interpreter

🌍World Economy

  • ▶The ripples can spread across borders, nudging growth forecasts here and there.
  • ▶Confidence among international firms might wobble until the picture clears.

🏙️Local Economy

  • ▶Small firms on your street might pass costs on carefully, bit by bit.
  • ▶Everyday essentials might nudge in price as suppliers adjust.

🏦Rates & Banks

  • ▶Borrowing plans are usually safe from sudden shocks over something like this.
  • ▶Interest rates and mortgage bills are unlikely to jump straight away from this alone.

❤️Health

  • ▶Worry has a way of spreading faster than the facts sometimes.
  • ▶A little perspective usually helps once the initial shock fades.

💷Wealth

  • ▶Retirement plans are seldom derailed by news of this size alone.
  • ▶Short-term wobbles like this tend to even out given enough time.

🏠Housing

  • ▶The property market tends to move slowly, so expect any change to take time.
  • ▶First-time buyers might keep half an eye on mortgage rates after this.
Share: 𝕏 Twitter Facebook LinkedIn WhatsApp

Editorial note: This analysis was produced by the News Effects Interpreter, an AI editorial tool that cross-references 1 independent news sources and contextualises events in terms of their real-world impact on ordinary people. Original reporting is linked above. News Effects does not alter the facts of source reports.