Function-Level Execution Feedback for Code Preference Optimization

Function-Level Execution Feedback for Code Preference Optimization

Process supervision has been shown to boost mathematical reasoning by naturally expressing intermediate steps as chains of thought, yet its application to code generation remains largely unexplored due to the absence of a standard definition of a “step.” To address this gap, researchers introduced STEP‑KTODER, a framework that treats each module‑level function within a decomposed multi‑function program as a distinct step and assigns binary correctness labels using automatically generated unit tests. By combining function‑level process supervision with outcome‑level feedback on the full program, STEP‑KTODER provides a code‑specific implementation of stepwise knowledge‑targeted optimization (KTO), enabling more precise preference learning for large language models (LLMs) tasked with coding.

Evaluations across several benchmark suites—including HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench—demonstrate that STEP‑KTODER consistently outperforms both outcome‑only KTO and direct preference optimization (DPO). A key finding from the analysis is that execution‑based labels are crucial: when LLMs act as judges to annotate step correctness, they tend to over‑predict function failures, corrupt positive step labels, and ultimately degrade the effectiveness of downstream preference optimization. By contrast, the automatically generated unit tests used in STEP‑KTODER yield reliable binary signals that guide the model toward generating correct functional components.

The release of the STEP‑KTODER code on GitHub invites the research community to build upon this approach, potentially extending process supervision to other programming contexts and refining step definitions. The work also underscores the broader importance of accurate execution‑based feedback in training LLMs for code generation, suggesting that future systems should prioritize concrete test‑driven supervision over heuristic judgments. As the framework integrates with existing arXivLabs infrastructure, it aligns with the platform’s commitment to openness, community collaboration, and data privacy, offering a tangible avenue for further advancements in automated code synthesis.

Sources cited: 📰 ArXiv AI ↗

⚡ Effects Interpreter

🌍World Economy

  • ▶Global commerce often shrugs off small shocks, but keeps one eye open.
  • ▶Cross-border money flows can subtly change direction after events like this.

🏙️Local Economy

  • ▶Everyday spending habits nearby may shift once the news sinks in.
  • ▶High street footfall and spending can shift subtly after news like this.

🏦Rates & Banks

  • ▶Base rate decisions are usually made on data trends, not single headlines.
  • ▶Savings rates can lag behind this kind of news by weeks, not days.

❤️Health

  • ▶Sleep and appetite can be the first quiet casualties of unsettling news.
  • ▶The strain, if any, tends to show up quietly in everyday life.

💷Wealth

  • ▶Your overall wealth picture is unlikely to be defined by a single story like this.
  • ▶Nest eggs can wobble briefly before finding their footing again.

🏠Housing

  • ▶Buyers and renters could notice only a mild drift, if anything at all.
  • ▶Anyone mid-purchase might want to keep an eye on how this unfolds.
Share: 𝕏 Twitter Facebook LinkedIn WhatsApp

Editorial note: This analysis was produced by the News Effects Interpreter, an AI editorial tool that cross-references 1 independent news sources and contextualises events in terms of their real-world impact on ordinary people. Original reporting is linked above. News Effects does not alter the facts of source reports.