{"id":1496852,"date":"2024-10-17T18:25:00","date_gmt":"2024-10-17T22:25:00","guid":{"rendered":"https:\/\/bugaluu.com\/news\/?p=1496852"},"modified":"2024-10-17T18:25:00","modified_gmt":"2024-10-17T22:25:00","slug":"nvidias-new-open-source-ai-model-beats-gpt-4o-on-benchmarks","status":"publish","type":"post","link":"https:\/\/bugaluu.com\/news\/nvidias-new-open-source-ai-model-beats-gpt-4o-on-benchmarks\/1496852\/","title":{"rendered":"Nvidia&#8217;s New Open-Source AI Model Beats GPT-4o On Benchmarks"},"content":{"rendered":"<p><span class=\"field field--name-title field--type-string field--label-hidden\">Nvidia&#8217;s New Open-Source AI Model Beats GPT-4o On Benchmarks<\/span><\/p>\n<div class=\"clearfix text-formatted field field--name-body field--type-text-with-summary field--label-hidden field__item\">\n<p><a href=\"https:\/\/cointelegraph.com\/news\/nvidia-open-source-ai-nemotron-surpasses-open-ai-gpt-4o\"><em>Authored by Tristan Greene via CoinTelegraph.com,<\/em><\/a><\/p>\n<p><strong>Nvidia unceremoniously launched a new artificial intelligence model on Oct 15 that\u2019s purported to outperform state-of-the-art AI systems including GPT-4o and Claude-3.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cms.zerohedge.com\/s3\/files\/inline-images\/01929aff-0a23-7791-98ee-adc2c64c.jpg?itok=1TFbQmCr\"><\/a><\/p>\n<p>According to a post on the X.com social media platform from the Nvidia AI Developer account, the new model, dubbed<strong> Llama-3.1-Nemotron-70B-Instruct, \u201cis a leading model\u201d on lmarena.AI\u2019s Chatbot Arena.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cms.zerohedge.com\/s3\/files\/inline-images\/01929b69-d07a-7967-a771-93eb74d1.jpg?itok=iohTmKHk\"><\/a><\/p>\n<p><em>Nvidia AI announces the benchmarks score for Nemotron. Source:\u00a0<\/em><a href=\"https:\/\/x.com\/NVIDIAAIDev\/status\/1846227767333212622\"><em>Nvidia AI<\/em><\/a><\/p>\n<h2>Nemotron<\/h2>\n<p>Llama-3.1-Nemotron-70B-Instruct is, essentially, a modified version of Meta\u2019s open-source Llama-3.1-70B-Instruct.<\/p>\n<p>The \u201cNemotron\u201d portion of the model\u2019s name encapsulates Nvidia\u2019s contribution to the end result.\u00a0<\/p>\n<p><strong>The Llama \u201cherd\u201d of AI models, as Meta refers to them, are meant to be used as open-source foundations for developers to build on.<\/strong><\/p>\n<p>In the case of Nemotron, Nvidia took up the challenge and developed a system designed to be more \u201chelpful\u201d than popular models such as OpenAI\u2019s ChatGPT and Anthropic\u2019s Claude-3.\u00a0<\/p>\n<p>Nvidia\u00a0<a href=\"https:\/\/build.nvidia.com\/nvidia\/llama-3_1-nemotron-70b-instruct\/modelcard\">used<\/a>\u00a0specially curated datasets, advanced fine-tuning methods, and its own state-of-the-art AI hardware to turn Meta\u2019s vanilla model into what might be the most \u201chelpful\u201d AI model on the planet.\u00a0<\/p>\n<p><a href=\"https:\/\/cms.zerohedge.com\/s3\/files\/inline-images\/01929b6c-1097-71dc-ac8d-2ab8792a.jpg?itok=rLerNLgE\"><\/a><\/p>\n<p><em>An engineer\u2019s post on X.com expressing excitement for Nemotron\u2019s capabilities. Source:\u00a0<\/em><a href=\"https:\/\/x.com\/ImSh4yy\/status\/1846589157721715086\"><em>Shayan Taslim<\/em><\/a><\/p>\n<p><em><strong>\u201cI asked it a few coding questions I usually ask to compare LLMs and got some of the best answers from this one. lol, holy shit.\u201d<\/strong><\/em><\/p>\n<h2>Benchmarking<\/h2>\n<p><strong>When it comes to determining which AI model is \u201cthe best,\u201d there\u2019s no clear-cut methodology.<\/strong> Unlike, for example, measuring the ambient temperature with a mercury thermometer, there isn\u2019t a single \u201ctruth\u201d that exists when it comes to AI model performance.\u00a0<\/p>\n<p>Developers and researchers have to determine how well an AI model performs the same as humans are evaluated: through comparative testing.\u00a0<\/p>\n<p>AI benchmarking involves giving different AI models the same queries, tasks, questions, or problems and then comparing the usefulness of the results. Often, due to the subjectivity of what is and isn\u2019t considered useful, human proctors are used to determine a machine\u2019s performance through blind evaluations.\u00a0<\/p>\n<p>In Nemotron\u2019s case, it appears that Nvidia is claiming the new model outperforms existing state-of-the-art models such as GPT-4o and Claude-3 by a fairly wide margin.<\/p>\n<p><a href=\"https:\/\/cms.zerohedge.com\/s3\/files\/inline-images\/01929b6c-a24a-7bdf-81a0-5f3abb84.jpg?itok=0ceEn1-2\"><\/a><\/p>\n<p><em>The top of the Chatbot Arena leaderboards. Source: LMArenea.AI<\/em><\/p>\n<p>The image above depicts the ratings on the automated \u201cHard\u201d test on the Chatbot Arena Leaderboards. While Nvidia\u2019s Llama-3.1-Nemotron-70B-Instruct doesn\u2019t appear to be listed anywhere on the boards, if the developer\u2019s claim that it scored an 85 on this test is valid, it would be the de facto top model in this particular section.\u00a0<\/p>\n<p><strong>What makes the achievement perhaps even more interesting is that Llama-3.1-70B is Meta\u2019s middle-tier open-source AI model. <\/strong><\/p>\n<p>There\u2019s a much larger version of Llama-3.1, the 405B version (where the number refers to how many billion parameters the model was tuned with).<\/p>\n<p><strong>By comparison, GPT-4o is\u00a0<a href=\"https:\/\/arxiv.org\/pdf\/2407.09519\">estimated<\/a>\u00a0to have been developed with over one trillion parameters.<\/strong><\/p>\n<\/div>\n<p>      <span class=\"field field--name-uid field--type-entity-reference field--label-hidden\"><a title=\"View user profile.\" href=\"https:\/\/cms.zerohedge.com\/users\/tyler-durden\" class=\"username\">Tyler Durden<\/a><\/span><br \/>\n<span class=\"field field--name-created field--type-created field--label-hidden\">Thu, 10\/17\/2024 &#8211; 14:25<\/span><\/p>\n<p>\u200b<a href=\"https:\/\/www.zerohedge.com\/technology\/nvidias-new-open-source-ai-model-beats-gpt-4o-benchmarks\" target=\"_blank\" class=\"\" rel=\"noopener\">https:\/\/www.zerohedge.com\/technology\/nvidias-new-open-source-ai-model-beats-gpt-4o-benchmarks<\/a>\u00a0<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Nvidia&#8217;s New Open-Source AI Model Beats GPT-4o On Benchmarks Authored by Tristan Greene via CoinTelegraph.com, Nvidia unceremoniously launched a new artificial intelligence model on Oct&#8230;<\/p>\n","protected":false},"author":0,"featured_media":1496853,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[1],"tags":[],"class_list":["post-1496852","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-news","wpcat-1-id"],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/pbimBl-6hoM","jetpack_featured_media_url":"https:\/\/bugaluu.com\/news\/wp-content\/uploads\/sites\/3\/2024\/10\/01929aff-0a23-7791-98ee-adc2c64c-3lx32f.jpeg","_links":{"self":[{"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/posts\/1496852","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/comments?post=1496852"}],"version-history":[{"count":0,"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/posts\/1496852\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/media\/1496853"}],"wp:attachment":[{"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/media?parent=1496852"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/categories?post=1496852"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/bugaluu.com\/news\/wp-json\/wp\/v2\/tags?post=1496852"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}