LLMs Have Different Taste on Coding

English 简体中文 繁体中文 ภาษาไทย Tiếng Việt
Summary

Large Language Models demonstrate distinct "tastes" or specializations when assisting with coding tasks, making model selection crucial for different development phases. Claude Sonnet 4 excels at scaffolding new projects and generating initial structures, but often over-engineers debugging efforts and struggles with refinement. Gemini 2.5 Pro proves effective for pinpointing and fixing specific bugs, though its tendency to aggressively refactor existing code can be counterproductive. In contrast, GPT-4.1 is ideal for incremental improvements and precise, small-scale changes, offering concise and reliable fixes without verbose explanations. Developers should ultimately choose an LLM whose style aligns with the specific task, potentially combining models for different project stages.

After experimenting with Cursor, WindSurf Pro, Gemini Cli, Claude Code and others, there is something interesting coming to mind that each LLM has its own taste when giving the answers/suggestions to coding question.

Here are some findings based on the using of the flagship models from each vendor.

Claude Sonnet 4 – The Trailblazer

This is the most popular model for coding so far. It usually does what you ask and will eventually get the job done as expected more or less comparing to other models.

Best for: Setting up new projects and scaffolding.

  • Instantly builds functional structure according to your requirements.
  • Excellent at defining routes, boilerplate, and high-level features.
  • It usually will help generate a working project or example and fulfill your requirement with the most update-to-date solutions available.

Weakness:

  • When debugging, it often misses the key issues. For example, it once misidentified a CSS class selector, not checking actual DOM elements but guessing based on experience.
  • Upon asking it to refine, it spiraled: added complex matching logic, nested click handlers, even tried clicking every nested element—really overengineered.
  • It tends to turn simple things complicated if keep asking it to refine some small things or rework on something which is not that expected.
  • It would produce lots of emojis or icons when providing responses and the code generated would also contain emojis(if adding debug messages in code), you could ask it to generate less emojis and it would follow. But it would do the same again after some time.

Once stuck in a loop, I switched away. Great at starting things—but not debugging them.

 Gemini 2.5 Pro – The Over-Refactorer

This model from Google seems respond very fast and can give suggestions on details but seems lacking bigger picture. It is more like catching and fixing model.

Best for: Fixing bugs.

If there is some bug in some specific area of the project, it would perform well on locating the issue and offering a fix. It once accurately addressed a click handling issue with a clear solution and it is immediately functional post the fix.

But:

  • For the next feature, it wrote verbose explanations and rewrote preceding code, essentially rolling back previous work.
  • It loves to refactor mid-flow, sometimes undoing earlier progress.
  • Oddly, after I scolded it, it apologized and fixed things—but the mental overhead felt exhausting.

It’s precise but emotionally draining to use.

GPT-4.1 – The Surgical Fixer

It is also a model which more or less can help complete the work but with more guidance and review needed. It tends to be very careful but not that confident at first glance.

Best for: Incremental improvements and precision.

  • Responds briefly and to the point.
  • Handles small changes perfectly—refines one function bit by bit.
  • About 70% successful on the first try; for the rest, responds calmly and quietly iterates.

Why I’ve stuck with it:

  • No long, convoluted explanations cluttering the agent window.
  • I found Gemini 2.5 explaining its greatness in code terms not helpful:
    1. I don’t understand high-tech bragging.
    2. If I do understand, I’d just write it myself.

I don’t need the model to pontificate—I just need things to work cleanly.

Comparing the Tastes

Model

Best For

Personality

Downside

Claude Sonnet 4

Scaffolding & bootstrapping

Fast, confident

Over-engineers complex tasks

Gemini 2.5 Pro

Bug fixes

Verbose, refactoring-focused

Aggressively rewrites prior code

GPT‑4.1

Small edits, reliable fixes

Precise, quiet

May need prompting for clarity

Final Thoughts: Choose Your Flavor

At the end of the day, LLMs have different coding tastes. It’s not about finding the smartest model—that’s overrated. It’s about finding the model whose style matches your workflow. That personal “taste” makes all the difference. Sometimes some models can be combined and used at the different stages of the same project. 

COMPARISON GUIDE OPENAI LLM GEMINI CLAUDE CODE TASTE GPT

  RELATED

  COMMENTS

0

No comment for this article.