[ad_1]
Last week, Meta unveiled Llama 2, a new large language model with up to 70 billion parameters. The new generative AI system is a spectacular blow across the bow of OpenAI, which shares few details about most AI models, including GPT-3/3.5 and GPT-4. The version of Llama 2 with 40 percent of ChatGPT parameters 3.5, according to Wikipedia, included a major partnership with Microsoft. And Redmond isn’t just a nominal partner, having recently announced support for Llama 2 on Azure and Windows. Meanwhile, Qualcomm now says they are entering the LLM fray with Llama 2 presentation plans to bring Llama 2 to smartphones.
Slightly more controversial is Metas’ claim that Llama 2 is open source. Meta and Microsoft certainly tout Llamas’ new open source credentials. (Some open source developers, on the other hand, beg to differ.)
The Llama 2 license is a force multiplier, giving developers and researchers the opportunity to fine-tune the model for their specific needs.
The developments of the last few weeks mean, whatever the origin, a dramatic expansion of the capacity and scope of open source AI models.
Oh, this is just so much better, says Aravind Srinivas, co-founder and CEO of Perplexity.ai. Whether or not they match GPT 3.5 1690221097It is only a matter of time.
Llama 2 – Optimized and ready to chat
Perplexity.ai offers an impressive free online demo of multiple Llama 2 templates. Its results are competitive with today’s best chatbots, including ChatGPT and Google Bard. Llama 2 quickly generates clean, natural text that, while unlikely to win awards, is easy to read and understand. Llama 2 can also output commonly understood facts, generate code, and solve mathematical equations.
Llama 2, like all LLMs, will occasionally generate incorrect or unusable answers, but the Metas paper introducing Llama 2 says it is on par with OpenAIs GPT 3.5 in academic benchmarks such as MMLU (which measures an LLM’s knowledge of 57 STEM subjects) and GSM8K (which measures an LLM’s understanding of mathematics).
Most of the small models that outstrip Llama 2 in the Open LLM rankings are themselves based on Meta’s predecessor model, Llama.
The Metas researchers achieved this in part thanks to the size of the model, but that’s only half the story. Llama 2 uses supervised tuning, reinforcement learning with human feedback, and a new technique called Ghost Attention (GAtt) which, according to Metas’ article, allows for multi-turn dialogue control. Put simply, GAtt helps Llama 2 generate the desired results when asked to work within a specific constraint, as might happen when asked to act as a historical figure or to produce answers in the context of a specific topic, such as architecture.
Llama 2s Ghost Attention, LLM promoters say, helps the model deliver conversational results that fit user-defined constraints. Half
These techniques help Llama 2 offer a wide range of models with solid benchmark performance relative to their size. The largest model, Llama 2 70B (with 70 billion parameters), performs best in all benchmarks, but Meta also supplies Llama 2 7B and Llama 2 13B.
Variants with fewer parameters don’t perform as well as the Llama 2 70B, but are compact enough to work locally on less powerful devices such as smartphones. Qualcomm, a leading system-on-chip (SoC) manufacturer for smartphones, announced a partnership with Meta to make Llama 2 work locally on Qualcomm-powered smartphones starting in 2024.
We can use our software tools to build and optimize the model specifically to run on our Hexagon processor, says Rodrigo Caruso Neves do Amaral, marketing communications specialist at Qualcomm. The amount of energy that is saved by running the device has a huge impact, both for the companies running these models and for the consumer who would sometimes have to pay to gain access to these applications.
Open source fits where closed models can’t
Running a large language model offline on a smartphone is something closed AI models (like OpenAI’s GPT 3.5 and Google’s PaLM2) can’t handle. This is not necessarily due to technical limitations (presumably, OpenAI and Google could offer a model suitable for a smartphone) but rather a philosophical division. OpenAI and Google provide LLM as an API. An internet connection is required to access the API, and customers pay based on usage.
Llama 2, by contrast, was released under a license that allows for unlimited, free commercial and academic use. The license does not meet all standards set forth by the Open Source Initiative, as the license includes a clause requiring permission to use Llama 2 for products or services with more than 700 million monthly active users. However, this clause is only relevant for Metas’ biggest competitors, such as OpenAI and Google. Metas Llama 2 models are already appearing in the HuggingFaces Open LLM leaderboard, with llama-2-70b-chat-hf claiming the third-best latency and throughput benchmark, as it closed on Monday 24 July. (AI developers are quickly tapping into the potential of Llamas 2: the current flagship model at press time, Stability AIs FreeWilly2, is actually already based on Llama 2, but FreeWilly2 fine-tunes the model with a different dataset.)
As of July 21, AI aggregator HuggingFaces OpenLLM Leaderboard has ranked llama-2-70b-chat-hf as the second best performer among all open LLMs in terms of performance and latency. Hugging Face
Srinivas sees Llama 2’s open source license as a force multiplier, giving developers and researchers the opportunity to fine-tune the model for their specific needs. One person can start a fork of Llama 2 where they focus on quantization, another person can start a fork of Llama where they focus on low rank tuning, [] another person can work on distilling larger patterns into smaller ones. Progress accelerates.
This will prove especially relevant for developers targeting cutting-edge devices such as smartphones. The fact that the Llama 2 70B performs well is no big surprise given the size of the model. But even the smaller Llama 2 models fare well in relation to the model size. And most of the small models that top Llama 2 in the Open LLM rankings are themselves based on Metas’ previous model, Llama. This suggests that Llama 2 will climb the ranks as developers from the open source community apply their talents to Llama 2.
I think that [Llama 2 7B and Llama 2 13B] they are already exciting. … This is just the beginning, right? [Meta] put it out and now people can improve it, says Srinivas. You can create other frameworks and other engineering layers and this gives more power to all.
From articles on your site
Related articles Around the web
|
Sources 2/ https://spectrum.ieee.org/llama-2-llm The mention sources can contact us to remove/changing this article |
[ad_2]