{"id":44113,"date":"2024-07-24T15:12:15","date_gmt":"2024-07-24T13:12:15","guid":{"rendered":"https:\/\/www.schneider.im\/?p=44113"},"modified":"2024-07-25T09:51:57","modified_gmt":"2024-07-25T07:51:57","slug":"microsoft-azure-ai-gpt-4o-mini-available","status":"publish","type":"post","link":"https:\/\/www.schneider.im\/lu\/microsoft-azure-ai-gpt-4o-mini-available\/","title":{"rendered":"Microsoft Azure AI: GPT-4o Mini available"},"content":{"rendered":"<p style=\"text-align: justify;\">Effective July 18, 2024, OpenAI\u2019s fastest model, <strong>GPT-4o mini<\/strong>, is <strong>available on Microsoft Azure OpenAI Service<\/strong>. This model offers significant <strong>improvements in speed, cost, and multilingual capabilities<\/strong>. It supports <strong>text processing with excellent speed<\/strong>, and <strong>image, audio, and video capabilities <\/strong>will be<strong> added later<\/strong>. Customers can <strong>try it for free<\/strong> in the <a href=\"https:\/\/oai.azure.com\/portal\/playground\">Azure OpenAI Studio Playground<\/a>.<img decoding=\"async\" class=\"size-full wp-image-44126 aligncenter\" src=\"https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-24-Update-Website-Microsoft-Azure-OpenAI-Service-GPT-4o.png\" alt=\"\" width=\"800\" height=\"514\" srcset=\"https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-24-Update-Website-Microsoft-Azure-OpenAI-Service-GPT-4o.png 800w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-24-Update-Website-Microsoft-Azure-OpenAI-Service-GPT-4o-300x193.png 300w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-24-Update-Website-Microsoft-Azure-OpenAI-Service-GPT-4o-150x96.png 150w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-24-Update-Website-Microsoft-Azure-OpenAI-Service-GPT-4o-768x493.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/p>\n<p>&nbsp;<\/p>\n<h2 style=\"text-align: justify;\"><strong>What is GPT-4o Mini?<\/strong><\/h2>\n<p style=\"text-align: justify;\"><strong>GPT-4o mini<\/strong> is a highly <strong>efficient AI model<\/strong> designed for <strong>fast and cost-effective application delivery<\/strong>. It is <strong>significantly smarter than GPT-3.5 Turbo<\/strong>, scoring 82% on the Measuring Massive Multitask Language Understanding (MMLU) benchmark compared to 70% for GPT-3.5 Turbo. It also offers a <strong>128K context window<\/strong> and <strong>improved multilingual capabilities<\/strong>. <strong>Fine-tuning<\/strong> for GPT-4o mini is available, allowing customers to <strong>customize the model for specific use cases and scenarios<\/strong>.<\/p>\n<p>&nbsp;<\/p>\n<h2 style=\"text-align: justify;\"><strong>Key Features<\/strong><\/h2>\n<ul style=\"text-align: justify;\">\n<li><strong>Speed and Cost<\/strong>: More than 60% less expensive than GPT-3.5 Turbo.<\/li>\n<li><strong>Performance<\/strong>: Scores 82% on MMLU vs. GPT-3.5 Turbo scoring only 70%.<\/li>\n<li><strong>Context Window<\/strong>: Expanded to 128K. The context window refers to the amount of text (measured in tokens) that the model can consider at once when generating responses. Essentially, it is the \u201cmemory\u201d of the model during a single interaction. For example, if a model has a context window of 16K tokens, it can take into account up to 16,000 tokens of text from the conversation history or input data to generate its response.<\/li>\n<li><strong>Multilingual Capabilities<\/strong>: Enhanced support for multiple languages.<\/li>\n<li><strong>Safety Features<\/strong>: Includes prompt shields that prevent the model from generating harmful or inappropriate content. It also has a protected material detection by default that ensures that the model does not share confidential content.<\/li>\n<li><strong>Data Residency<\/strong>: Available in 27 regions, including 9 regions in Europe. Find an up-to-date list here: <a href=\"https:\/\/go.microsoft.com\/fwlink\/?linkid=2274842&amp;clcid=0x409\">https:\/\/go.microsoft.com\/fwlink\/?linkid=2274842&amp;clcid=0x409<\/a>.<\/li>\n<li><strong>Global Pay-As-You-Go<\/strong>: Flexible payment options with a high throughput limit of 15M tokens per minute (TPM).<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<p style=\"text-align: justify;\"><strong><img decoding=\"async\" class=\"size-full wp-image-44114 alignright\" src=\"https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Microsoft-Azure-OpenAI-Service.jpg\" alt=\"\" width=\"225\" height=\"225\" srcset=\"https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Microsoft-Azure-OpenAI-Service.jpg 225w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Microsoft-Azure-OpenAI-Service-150x150.jpg 150w\" sizes=\"(max-width: 225px) 100vw, 225px\" \/><\/strong><\/p>\n<h2 style=\"text-align: justify;\"><strong>Licensing<\/strong><\/h2>\n<p style=\"text-align: justify;\"><strong>GPT-4o mini<\/strong> is available under <strong>Azure AI\u2019s global pay-as-you-go<\/strong> deployment at <strong>0.15$ per million input tokens*<\/strong> and <strong>0.60$ per million output tokens*<\/strong>. This model is, like GPT-3.5 Turbo and GPT-4o, also available on <strong>Azure AI\u2019s Batch service<\/strong>, offering <strong>high throughput jobs at a discounted rate<\/strong>. <strong>Batch<\/strong> <strong>delivers<\/strong> high throughput jobs <strong>within 24 hours of submission at a 50% discount rate<\/strong> by using off-peak capacity. Off-peak capacity refers to times when the demand for computational resources is lower. By utilizing these periods, the service can offer a discount because the resources are less in demand and therefore cheaper to use.<\/p>\n<p>&nbsp;<\/p>\n<h2 style=\"text-align: justify;\"><strong>Comparing models<\/strong><\/h2>\n<table width=\"990\">\n<tbody>\n<tr>\n<td width=\"247\"><strong>Feature<\/strong><\/td>\n<td width=\"247\"><strong>GPT-4o Mini<\/strong><\/td>\n<td width=\"248\"><strong>GPT-4o<\/strong><\/td>\n<td width=\"248\"><strong>GPT-3.5 Turbo<\/strong><\/td>\n<\/tr>\n<tr>\n<td width=\"247\"><strong>Quality<\/strong> Index<\/td>\n<td width=\"247\">85<\/td>\n<td width=\"248\">100<\/td>\n<td width=\"248\">59<\/td>\n<\/tr>\n<tr>\n<td width=\"247\"><strong>MMLU<\/strong> Score<\/td>\n<td width=\"247\">82%<\/td>\n<td width=\"248\">88.7%<\/td>\n<td width=\"248\">70%<\/td>\n<\/tr>\n<tr>\n<td width=\"247\"><strong>Context<\/strong> <strong>Window<\/strong><\/td>\n<td width=\"247\">128K tokens<\/td>\n<td width=\"248\">128K tokens<\/td>\n<td width=\"248\">16K tokens<\/td>\n<\/tr>\n<tr>\n<td width=\"247\"><strong>Speed\u00a0<\/strong>(Output Tokens per Second)<\/td>\n<td width=\"247\">108 tokens\/sec<\/td>\n<td width=\"248\">83 tokens\/sec<\/td>\n<td width=\"248\">79 tokens\/sec<\/td>\n<\/tr>\n<tr>\n<td width=\"247\"><strong>Latency\u00a0<\/strong>(Seconds to First Tokens Chunk Received; Lower is better)<\/td>\n<td width=\"247\">0.53<\/td>\n<td width=\"248\">0.44<\/td>\n<td width=\"248\">0.37<\/td>\n<\/tr>\n<tr>\n<td width=\"247\"><strong>Modalities<\/strong> Supported<\/td>\n<td width=\"247\">Text, (future: Image, Audio, Video)<\/td>\n<td width=\"248\">Text, Vision, Audio, Video<\/td>\n<td width=\"248\">Text<\/td>\n<\/tr>\n<tr>\n<td width=\"247\"><strong>Availability<\/strong><\/td>\n<td width=\"247\">27 regions<\/td>\n<td width=\"248\">27 regions<\/td>\n<td width=\"248\">27 regions<\/td>\n<\/tr>\n<tr>\n<td width=\"247\">Standard Pricing (<strong>Input<\/strong> Tokens)<\/td>\n<td width=\"247\">$0.15 \/ 1M tokens*<\/td>\n<td width=\"248\">$5.00 \/ 1M tokens*<\/td>\n<td width=\"248\">$0.50 \/ 1M tokens*<\/td>\n<\/tr>\n<tr>\n<td width=\"247\">Standard Pricing (<strong>Output<\/strong> Tokens)<\/td>\n<td width=\"247\">$0.60 \/ 1M tokens*<\/td>\n<td width=\"248\">$15.00 \/ 1M tokens*<\/td>\n<td width=\"248\">$1.50 \/ 1M tokens*<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p style=\"text-align: justify;\">*Prices may change.<\/p>\n<p>&nbsp;<\/p>\n<h2 style=\"text-align: justify;\"><strong>Performance Comparison<\/strong><\/h2>\n<p style=\"text-align: justify;\"><img decoding=\"async\" class=\"aligncenter wp-image-44118 size-full\" src=\"https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Website-Microsoft-Azure-OpenAI-Service-Comparing-models.png\" alt=\"\" width=\"828\" height=\"438\" srcset=\"https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Website-Microsoft-Azure-OpenAI-Service-Comparing-models.png 828w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Website-Microsoft-Azure-OpenAI-Service-Comparing-models-300x159.png 300w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Website-Microsoft-Azure-OpenAI-Service-Comparing-models-800x423.png 800w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Website-Microsoft-Azure-OpenAI-Service-Comparing-models-150x79.png 150w, https:\/\/www.schneider.im\/media\/2024\/07\/SCHNEIDER-IT-MANAGEMENT-2024-07-22-Update-Website-Microsoft-Azure-OpenAI-Service-Comparing-models-768x406.png 768w\" sizes=\"(max-width: 828px) 100vw, 828px\" \/><\/p>\n<p>&nbsp;<\/p>\n<h2 style=\"text-align: justify;\"><strong>More information<\/strong><\/h2>\n<p style=\"text-align: justify;\">Find the <strong>announcement<\/strong> here: <a href=\"https:\/\/azure.microsoft.com\/en-us\/blog\/openais-fastest-model-gpt-4o-mini-is-now-available-on-azure-ai\/\">https:\/\/azure.microsoft.com\/en-us\/blog\/openais-fastest-model-gpt-4o-mini-is-now-available-on-azure-ai\/<\/a>.<\/p>\n<p style=\"text-align: justify;\">Find more <strong>pricing info<\/strong> here &#8211; OpenAI&#8217;s pricing is the same as the pricing in Azure OpenAI Studio: <a href=\"https:\/\/openai.com\/api\/pricing\/\">https:\/\/openai.com\/api\/pricing\/<\/a>.<\/p>\n<p style=\"text-align: justify;\">Find a <strong>globe of Microsoft Datacenters<\/strong> here: <a href=\"https:\/\/datacenters.microsoft.com\/globe\/explore\/\">https:\/\/datacenters.microsoft.com\/globe\/explore\/<\/a>.<\/p>\n<p style=\"text-align: justify;\">Find an interactive model comparison here: <a href=\"https:\/\/artificialanalysis.ai\/models\/gpt-4o-mini\/providers\">https:\/\/artificialanalysis.ai\/models\/gpt-4o-mini\/providers<\/a>.<\/p>\n<p style=\"text-align: justify;\">For more on <strong>Microsoft licensing<\/strong>, visit <strong>our<\/strong> <strong>Microsoft vendor<\/strong> <strong>page<\/strong> at: <a href=\"https:\/\/www.schneider.im\/software\/microsoft\/\">https:\/\/www.schneider.im\/software\/microsoft\/<\/a>.<\/p>\n<p style=\"text-align: justify;\">Please <a href=\"https:\/\/www.schneider.im\/contact\/\">contact us<\/a> for <strong>expert services<\/strong> on your specific <strong>Microsoft<\/strong> software and Online Services requirements and to <strong>request a quote today<\/strong>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Effective July 18, 2024, OpenAI\u2019s fastest model, GPT-4o mini, is available on Microsoft Azure OpenAI Service. This model offers significant improvements in speed, cost, and multilingual capabilities. It supports text processing with excellent speed, and image, audio, and video capabilities will be added later. Customers can try it for free in the Azure OpenAI Studio [&hellip;]<\/p>\n","protected":false},"author":11,"featured_media":33016,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[248],"tags":[],"class_list":["post-44113","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-microsoft-azure"],"_links":{"self":[{"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/posts\/44113","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/users\/11"}],"replies":[{"embeddable":true,"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/comments?post=44113"}],"version-history":[{"count":5,"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/posts\/44113\/revisions"}],"predecessor-version":[{"id":44141,"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/posts\/44113\/revisions\/44141"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/media\/33016"}],"wp:attachment":[{"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/media?parent=44113"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/categories?post=44113"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.schneider.im\/lu\/wp-json\/wp\/v2\/tags?post=44113"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}