HomeApp PortfolioAI Case StudiesBlogSearch
AI & Agentic AI
AI Consulting and Implementation Gen AI Chatbot Agentic AI with N8N Anthropic Claude Partner Malaysia: AI Development for Enterprise By Claude Certified Architect (CCA-F) Claude Select Partner Malaysia and Singapore Claude AI Training Malaysia (Technical): CCAR-F Preparation Class Claude AI Training Malaysia (Non-Technical): CCAO-F Preparation Class OpenAI ChatGPT Partner Malaysia: Enterprise AI Solutions Gen AI OPC One Person Company AI Forward-Deployed Engineers FDE as a Service Enterprise Voice AI Agent GPU-as-a-Service (GPUaaS) in Malaysia Malaysia GPU as a Service Guides: From GPU, DC, Power, Software, KYC to Pricing AI Gateway Malaysia | LiteLLM, Bifrost, Quota Management and Enterprise AI Governance AI and Data for General Elections ReelHero AI: AI Templated Video Generator Gen AI Digital Avatar AI Digital Banking Malaysia Conversational Payment Solutions AI Chatbot Automated Testing: LLM as a Judge and Red-Teaming Malaysia AI Vibe Coding/VibeOps LLM Fine-Tuning as a Service MySTI-Certified AI Tech Products AI Deepfake Detection
Development & Apps
Our Services AR/VR/XR Apple Vision Pro Development Blockchain/Web3/Smart Contract UI/UX Design Backend Development Power BI Dashboard EV Solutions – OCPP/OCPI/E-MSP Web App Strapi Headless CMS Development SuperApp Development Flutter React Native Native iOS Native Android CarPlay/Android Auto In Car Assistant Project Management Enterprise Grade Test Automation For Software Applications Coding Training Outsystems
Cloud, Fintech & Enterprise
AWS Malaysia Cloud Modernization: Your Trusted AWS Partner Use Case: JomeInvoice Use Case: DahReply Azure Malaysia Cloud Modernization: Your Trusted Microsoft Partner Malaysia Cloud Modernization with Local BytePlus Cloud: Your Trusted BytePlus Partner Lark Malaysia Reseller For SME and Corporates Whitelabel E-Wallet Development Regulated Fintech Development Clean Energy Digital Solutions Malaysia Malaysia IOT Platform/Hardware Integration Software Development MyDigital ID Integration Property Development Digital Platforms Malaysia Cloud Framework Agreement CFA for Government Enterprise Grade IRB/LHDN E-Invoicing Middleware Meta WhatsApp and SMS API Services Government Grant
Team Extension / Staff Augmentation
Team Extension Java Springboot Developer Staff Augmentation and Outsourcing Rust Developer Staff Augmentation and Outsourcing Golang Developer Staff Augmentation and Outsourcing C# .NET Core Developer Staff Augmentation and Outsourcing MuleSoft Developer Staff Augmentation and Outsourcing Gen AI LLM Developer Staff Augmentation
About Us
Our Standards – CMMI and ISOs Certified Life at Agmo Careers – Join Agmoians Agmo Singapore Agmo Academy
Request a Quote/Demo
AI

GPT-5: Breaking Through the Limits of Large Language Models

by Tan Aik Keong (AK)

Large language models have driven a wave of change across tech in the past few years — from ChatGPT to its many derivative applications, becoming a core part of the productivity toolkit. But earlier models still carried several key problems: hallucination (fabricated information), sycophancy, opaque failure modes, and safety and ethics risk. The newly released GPT-5 delivers real progress on exactly these pain points, moving toward a more reliable, more controllable system.

Model routing

Older models often took a one-size-fits-all approach, handling a simple question and a complex reasoning task the same way. GPT-5's internal model-routing mechanism calls different sub-models for different tasks — gpt-5-main and gpt-5-thinking, for instance — using the deeper version for demanding reasoning and the faster version for everyday Q&A, balancing performance and accuracy.

Reducing the hallucination rate

Hallucination is when a model generates plausible-sounding but incorrect information in the absence of a factual basis. Per OpenAI's system card, GPT-5 significantly reduces hallucination rates versus GPT-4 across multiple benchmarks, particularly on medical Q&A tasks like HealthBench Hard. More importantly, when it can't guarantee a correct answer, it practises failure transparency — openly admitting it can't answer, rather than making something up.

Reducing sycophancy

Earlier models often agreed blindly with whatever a user asserted, a sycophancy bias. GPT-5 introduces adversarial training and data refinement to reduce that tendency. Faced with a biased or factually wrong question, it holds a factual position rather than simply telling the user what they want to hear — particularly important on social or policy-sensitive topics.

Safe completions

Traditional safety mechanisms often rely on outright refusal, at the cost of user experience. GPT-5 uses a safe-completions approach, providing as much useful information as it safely can rather than shutting the question down. In medical or financial queries, for example, it offers a safe, filtered reference rather than a flat refusal — balancing safety with practical usefulness, and making the model more usable in professional settings.

Less deceptive output

Deceptive output is when a model pretends to have succeeded at a task it isn't actually capable of. GPT-5 shows clear improvement here too — per the system card, its rate of deceptive responses has dropped significantly, and it's more willing to state plainly that information is insufficient, which builds real user trust in the system.

Multimodal understanding

While GPT-5's core improvements are concentrated in text reasoning, it retains image input support, showing higher accuracy analysing complex charts and combined visual-and-text scenarios. Future research may extend further into voice and video, but publicly available material currently confirms strengthened image-text dual-modality specifically.

Controllability and transparency

GPT-5 also makes progress on output controllability. Users can use parameterised instructions to control the depth, style and sourcing of a response — a rigorous, citation-backed answer for an academic context, or a more conversational tone for everyday use. That controllability makes the model feel more like a customisable assistant, and less like a black box.

From demo to infrastructure

Overall, GPT-5's improvements aren't a single point upgrade — they're a systematic correction of LLMs' known limitations: from reducing hallucination rates to safe completions, from mitigating sycophancy to reducing deception. Together, they're moving AI from a stage performer into something you can actually depend on: infrastructure. As multimodality and controllability mature further, we may genuinely be entering a new stage of trustworthy AI.


Part of the AK AI Corner column. Originally published in Oriental Daily (东方日报) on Sep 20, 2025.