HN 日本語サマリー

← 一覧へ戻る
AI・機械学習

GLM-5.3 人工分析ベンチマーク

GLM-5.3 Artificial Analysis Benchmarks (artificialanalysis.ai)

105 pointsby apitman43 コメント

要約

GLM-5.3 (max) は、人工分析のインテリジェンス指数で8位/181モデル中という高い評価を得ており、同価格帯のモデルと比較しても競争力のある価格設定です。1Mトークンのコンテキストウィンドウをサポートし、テキスト入出力に対応しています。ただし、速度に関するデータは提供されていません。

全文翻訳

Artificial AnalysisKZ AI•Proprietary model•Released August 2026GLM-5.3 (max) Intelligence, Performance & Price AnalysisCompareAPI Provider Benchmarks Model summaryIntelligence#8 / 18160Artificial Analysis Intelligence Index4 out of 4 units for Intelligence.SpeedN/AOutput tokens per secondUnknown out of 4 units for Speed.Cost#53 / 181In $1.40Out $4.40Cache Discount 81%$0.68Cost per Intelligence Index task3 out of 4 units for Cost.Verbosity#72 / 181170MOutput tokens from Intelligence Index4 out of 4 units for Verbosity.Comparison SummaryGLM-5.3 (max) is amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price. The model supports text input, outputs text, and has a 1M tokens context window.GLM-5.3 (max) scores 60 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 35). When evaluating the Intelligence Index, it generated 170M tokens, which is very verbose in comparison to the median of 72M.Pricing for GLM-5.3 (max) is $1.40 per 1M input tokens (moderately priced, median: $1.75) and $4.40 per 1M output tokens (moderately priced, median: $10.00). In total, it cost $1238.50 to evaluate GLM-5.3 (max) on the Intelligence Index.Technical specificationsReasoningYesThis page shows the reasoning version of this model.A non-reasoning variant may also exist.Input modalitySupports: textOutput modalitySupports: textContext window1M~1500 A4 pages of size 12 Arial font181 models in this classMetrics are compared against models of the same class:Non-reasoning models → compared only with other non-reasoning modelsReasoning models → compared across both reasoning and non-reasoningOpen weights models → compared only with other open weights models of the same size class:Tiny: ≤4B parametersSmall: 4B–40B parametersMedium: 40B–150B parametersLarge: >150B parametersProprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:<$0.15 per 1M tokens$0.15–$1 per 1M tokens>$1 per 1M tokensHighlightsIntelligenceArtificial Analysis Intelligence Index · Higher is betterSpeedOutput tokens per second · Higher is betterCost per TaskWeighted average cost (USD) per Intelligence Index task · Lower is betterPrompt OptionsIntelligenceArtificial Analysis Intelligence IndexUpdatedAgentic IndexUpdatedArtificial Analysis Intelligence IndexArtificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR29 of 610 modelsNEWAdd model from specific providerReasoning models are indicated by a lightbulb iconArtificial Analysis Intelligence IndexArtificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.Open Weights / ProprietaryReasoning / Non-ReasoningText Only / Multimodal InputsArtificial Analysis Intelligence Index by Open Weights / ProprietaryArtificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR29 of 610 modelsNEWAdd model from specific providerProprietaryOpen Weights (Commercial Use Restricted)Open WeightsReasoning models are indicated by a lightbulb iconArtificial Analysis Intelligence IndexArtificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.Open WeightsIndicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license).BenchmarksIntelligence EvaluationsIntelligence evaluations measured independently by Artificial Analysis · Higher is betterCodingTool UseLong ContextMultimodalInstruction FollowingFaithfulnessWritingUser InteractionBusinessFinanceLegalMedicalSee more19 of 23 evaluations29 of 610 modelsNEWAdd model from specific providerGDPval-AA v2Agentic real-world work tasks, (Elo-500)/2000𝜏³-BankingUpdatedAgentic tool useTerminal-Bench v2.1Agentic coding & terminal useSciCodeCodingHumanity's Last ExamUpdatedReasoning & knowledgeGPQA DiamondScientific reasoningCritPtPhysics reasoningAA-Omniscience AccuracyUpdatedKnowledgeAA-Omniscience Non-Hallucination RateUpdated1 - hallucination rateAA-LCRUpdatedLong context reasoningAA-BriefcaseAgentic knowledge work, EloAutomationBench-AAAgentic SaaS workflowsHarvey LAB-AALegal agentic work, criterion pass rateEnterpriseOps-Gym-AAAgentic business operationsAA-AnalystAgentNewQuantitative analysis on spreadsheets & documentsIFBenchInstruction followingAPEX-Agents-AALong-horizon agentic tasksITBench-AAKubernetes incident root-cause analysisMMMU-ProVisual reasoningReasoning models are indicated by a lightbulb iconIntelligence Evaluation RelevanceWhile model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.Artificial Analysis Intelligence IndexArtificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.AA-OmniscienceAA-Omniscience IndexAA-Omniscience AccuracyAA-Omniscience Hallucination RateAA-Omniscience IndexAA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.29 of 481 modelsNEWAdd model from specific providerReasoning models are indicated by a lightbulb iconAA-Omniscience IndexAA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.Intelligence Index ComparisonsIntelligence Index vs. Cost per TaskIntelligence Index vs. Time per TaskIntelligence Index vs. Output SpeedIntelligence Index vs. End-to-End Response TimeIntelligence Index vs. Cost per Intelligence Index TaskArtificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task29 of 610 modelsNEWMost attractive quadrantPareto lineZ AIGoogleAnthropicDeepSeekSpaceXAIKimiNVIDIAMetaOpenAIAlibabaReasoning models are indicated by a lightbulb iconCost per Intelligence Index TaskWeighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.Artificial Analysis Intelligence IndexArtificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.Token UseOutput Tokens per TaskIntelligence Index vs. Output Tokens per TaskIntelligence Index Token UseIntelligence Index vs. Token UseOutput Tokens per Intelligence Index TaskWeighted average number of output tokens used to run