by Tan Aik Keong (AK)
Writing a column from Kuala Lumpur, watching Silicon Valley's giants trade blows over tooling always has a bit of a familiar ring to it. Salesforce CEO Marc Benioff, who says he's used ChatGPT daily for three straight years, spent two hours with Google's new Gemini 3 and then publicly declared on social media, "I'm not going back" — siding openly with Gemini 3. That moment pushed the Gemini 3 vs ChatGPT 5.1 showdown into headlines worldwide.
But for readers here, the question was never "who wins outright" — it's which one is actually strong at what, and what should you actually use each one for? Setting the drama aside, it's worth comparing Gemini 3 and ChatGPT 5.1 as two differently styled "smart colleagues."
Two different personalities
Gemini 3 behaves more like a multimodal reasoning engine. Google built it from the start to "understand everything" — text, images, audio, video, PDFs, code, all droppable into the same conversation, for it to trace the thread and pull out the key points. It's best suited to the "desk covered in papers" scenario — a stack of reports, a stack of charts, a stack of screenshots, and one question: "help me pull out the conclusions and the action plan." It works hard to assemble that whole picture.
ChatGPT 5.1 behaves more like an on-call chief of staff — the member of the GPT-5 family specifically tuned for conversational experience. Inside ChatGPT it runs two modes, Instant and Thinking: the former for quick back-and-forth, the latter for complex, multi-step problems. The system automatically switches between the two based on how hard the question is — in plain terms, "think harder when it matters, don't ramble when it doesn't." Its overall character is built for long conversations — editing a draft, writing code, working through a plan — and it feels comparatively natural and smooth to use over time.
Context length and multimodal depth
The real gap shows up in context length and multimodal depth. Gemini 3 Pro's input window runs to roughly a million tokens, with output in the tens of thousands — meaning it can absorb several times more information than the previous generation. Paired with native multimodality, you can feed it hours of transcribed meeting audio, a batch of design files, even an entire codebase in one go, and have it pull out the architecture, flag problems, and draft a refactoring plan. For messy, large-volume data, that combination has a real edge.
ChatGPT 5.1 takes a different approach here. Its Thinking mode offers a context window of around 190,000-plus tokens — considerably roomier than a typical model, but not chasing the million-token mark deliberately. OpenAI's strategy is a medium-to-large context window combined with conversation-history compression and caching to thread multi-turn conversations together, plus a memory feature so the model remembers a user's preferences and past tasks. The result: it can't swallow an entire "mountain of data" in one sitting, but in an everyday workflow it's very good at keeping cause and effect connected smoothly — a strong long-term working companion.
Automation and agent capability
The two also emphasise different things on automation and agent capability. Google has positioned Gemini 3 as the "heart" of various agentic workflows, particularly for reading legacy codebases, generating migration scripts, writing tests automatically, and operating internal tools. Paired with their new developer platform, it can be packaged as a kind of digital engineer — understanding your system first, then planning its own steps, calling tools, and taking on a chunk of the tedious work itself.
ChatGPT 5.1 can drive complex agent systems too, but its selling point is tunability and control. Developers can set different reasoning intensity levels, giving the model different degrees of caution when making API calls, editing code, or issuing commands — and for industries operating close to compliance boundaries, like finance or healthcare, that kind of adjustable safety margin often matters more than raw intelligence. You could say Gemini 3 is more like a capable operator who charges ahead on its own, while ChatGPT 5.1 is more like a control console that manages the pace.
What actually matters for you
For a Malaysian reader, the final, practical question is which ecosystem you're already in. If your business is already deep into Google's stack — Android, Chrome, Workspace, GCP — making Gemini 3 your default AI engine has the lowest integration cost; a lot of everyday work, like organising documents, analysing spreadsheets, generating a front-end prototype, can happen directly inside Google's own products. Conversely, if your infrastructure sits on the Microsoft-and-OpenAI side — Azure, Office, tools built around ChatGPT — keeping ChatGPT 5.1 as your primary model saves a lot of migration hassle.
Benioff can afford to declare "I'm not going back" — that's a personal call from someone running a giant company. For most businesses and everyday users, the more realistic move probably isn't picking a side at all — it's learning to run both: hand the heavy lifting to Gemini 3 when you're dealing with massive, multimodal, genuinely complex data, and keep ChatGPT 5.1 close by for long-running conversations, steady writing help, thinking things through, and everyday automation. Rather than asking who wins, it's more useful to work out which kind of "smart" actually helps you get today's work done.
Part of the AK AI Corner column. Originally published in Oriental Daily (东方日报) on Nov 29, 2025.
