Deep learning is a type of artificial intelligence that uses layered neural networks to analyze data. It can be used to process large amounts of data, including numbers, images, and even audio. Learn more about deep learning by visiting the website.

seen from United Kingdom

seen from United States
seen from Bulgaria

seen from Malaysia

seen from United Kingdom
seen from Brazil

seen from Malaysia

seen from Bulgaria
seen from Peru

seen from United States
seen from Philippines
seen from China
seen from United States
seen from China
seen from Kazakhstan

seen from France
seen from Belgium
seen from United States

seen from United States
seen from China
Deep learning is a type of artificial intelligence that uses layered neural networks to analyze data. It can be used to process large amounts of data, including numbers, images, and even audio. Learn more about deep learning by visiting the website.
Memory-Efficient Training on Habana® Gaudi® with DeepSpeed
One of the key challenges in Large Language Model (LLM) training is reducing the memory requirements needed for training without sacrificing compute/communication efficiency and model accuracy. DeepSpeed [2] is a popular deep learning software library which facilitates memory-efficient training of large language models. DeepSpeed includes ZeRO (Zero Redundancy Optimizer), a memory-efficient approach for distributed training [5]. ZeRO has multiple stages of memory efficient optimizations, and Habana’s SynapseAI® software currently supports ZeRO-1 and ZeRO-2. In this article, we will talk about what ZeRO is and how it is useful for training LLMs. We will provide a brief technical overview of ZeRO, covering ZeRO-1 and ZeRO-2 stages of memory optimization. More details on DeepSpeed Support on Habana SynapseAI Software can be found at Habana DeepSpeed User Guide. Now, let us dive into why we need memory efficient training for LLMs and how ZeRO can help achieve this.
Emergence of Large Language Models
Large Language Models (LLMs) are becoming super large, with model sizes growing by 10x in only a few years as shown in Figure 1 [7]. Increase in model sizes offers considerable gains in model accuracy. Large LLMs such as GPT-2 (1.5B), Megatron-LM (8.3B), T5 (11B), Turing-NLG (17B), Chinchilla (70B), GPT-3 (175B), OPT-175B, BLOOM (176B), etc. have been released to excel in various tasks such as natural language understanding, question answering, summarization, translation, and natural language generation. As the size of LLMs keeps growing, how can we efficiently train such large models? Of course, the answer is “parallelization”