{"id":860,"date":"2023-10-05T09:04:40","date_gmt":"2023-10-05T09:04:40","guid":{"rendered":"https:\/\/tbekk.com\/devstream\/?p=860"},"modified":"2023-10-05T09:50:01","modified_gmt":"2023-10-05T09:50:01","slug":"weekly-ai-and-nlp-news-october-2nd-2023","status":"publish","type":"post","link":"https:\/\/tbekk.com\/devstream\/2023\/10\/05\/weekly-ai-and-nlp-news-october-2nd-2023\/","title":{"rendered":"Weekly AI and NLP News \u2014 October 2nd 2023"},"content":{"rendered":"\n<h2 class=\"wp-block-heading has-medium-gray-color has-text-color\">Voice and image capabilities into ChatGPT, Amazon invests 4B$ in Anthropic, and Mistral LLM<\/h2>\n\n\n\n<hr class=\"wp-block-separator has-text-color has-light-gray-color has-alpha-channel-opacity has-light-gray-background-color has-background is-style-wide\"\/>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em><strong>Link:<\/strong><\/em> <a href=\"https:\/\/medium.com\/nlplanet\/weekly-ai-and-nlp-news-october-2nd-2023-5b8bcbd62721\"><em>NLPlanet<\/em><\/a><\/li>\n\n\n\n<li><strong><em>Author:<\/em><\/strong> <a href=\"https:\/\/medium.com\/@chiusanofabio94?source=post_page-----3e128fbda17d--------------------------------\"><em>Fabio Chiusano<\/em><\/a><\/li>\n\n\n\n<li><strong><em>Publication Date:<\/em><\/strong> <em>Oct. 2, 2023<\/em><\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-text-color has-light-gray-color has-alpha-channel-opacity has-light-gray-background-color has-background is-style-wide\"\/>\n\n\n\n<p id=\"cb1a\">Here are your weekly articles, guides, and news about NLP and AI chosen for you by\u00a0<a rel=\"noreferrer noopener\" href=\"https:\/\/www.nlplanet.org\/\" target=\"_blank\">NLPlanet<\/a>!<\/p>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"7727\">\ud83d\ude0e News From The Web<\/h1>\n\n\n\n<ul class=\"wp-block-list\">\n<li><a href=\"https:\/\/openai.com\/blog\/chatgpt-can-now-see-hear-and-speak\" rel=\"noreferrer noopener\" target=\"_blank\">ChatGPT can now see, hear, and speak<\/a>. OpenAI has introduced new voice and image capabilities to their AI assistant, ChatGPT. Users can now engage in natural voice conversations and receive relevant responses. The image feature enables users to show ChatGPT images for assistance in interpretation.<\/li>\n\n\n\n<li><a href=\"https:\/\/www.anthropic.com\/index\/anthropic-amazon\" rel=\"noreferrer noopener\" target=\"_blank\">Amazon will invest up to $4 billion in Anthropic<\/a>. Amazon has made a significant $4 billion investment in Anthropic. This partnership will enable Anthropic to benefit from Amazon Web Services (AWS), specifically leveraging AWS\u2019s Trainium and Inferentia chips to enhance model training and deployment capabilities. Additionally, Anthropic will use Amazon Bedrock to optimize Claude versions and explore finetuning options.<\/li>\n\n\n\n<li><a href=\"https:\/\/mistral.ai\/news\/announcing-mistral-7b\/\" rel=\"noreferrer noopener\" target=\"_blank\">Mistral 7B<\/a>. The Mistral 7B model, powered by Grouped-query attention (GQA) and Sliding Window Attention (SWA), outperforms other models in various domains while maintaining strong performance in both English and coding tasks. Its impressive benchmarks make it the best 7B model to date, enhancing inference speed and sequence handling efficiency.<\/li>\n\n\n\n<li><a href=\"https:\/\/www.crn.com\/news\/components-peripherals\/llm-startup-embraces-amd-gpus-says-rocm-has-parity-with-nvidia-s-cuda-platform\" rel=\"noreferrer noopener\" target=\"_blank\">LLM Startup Embraces AMD GPUs, Says ROCm Has \u2018Parity\u2019 With Nvidia\u2019s CUDA Platform<\/a>. A startup called Lamini is using over 100 AMD Instinct MI200 GPUs and found that AMD\u2019s ROCm software platform rivals Nvidia\u2019s CUDA platform. They claim that running a large language model on their platform is 10x cheaper than on Amazon Web Services.<\/li>\n\n\n\n<li><a href=\"https:\/\/decrypt.co\/198987\/openai-brings-web-search-back-to-chatgpt\" rel=\"noreferrer noopener\" target=\"_blank\">OpenAI\u2019s ChatGPT Now Searches the Web in Real Time \u2014 Again<\/a>. OpenAI has reintroduced web searching for ChatGPT, allowing users to access recent information. Important updates have been made, including compliance with robots.txt rules and user agent identification, giving websites more control. Currently available to Plus and Enterprise users, expansion plans are in progress.<\/li>\n<\/ul>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"66b4\">\ud83d\udcda Guides From The Web<\/h1>\n\n\n\n<ul class=\"wp-block-list\">\n<li><a href=\"https:\/\/blog.roboflow.com\/gpt-4-vision\/\" rel=\"noreferrer noopener\" target=\"_blank\">First Impressions with GPT-4V(ision)<\/a>. OpenAI has released GPT-4V for Plus users, showcasing its image processing skills, OCR capabilities, and performance in solving mathematical problems. However, it still faces challenges with object detection and struggles with CAPTCHA and grid-based puzzles.<\/li>\n\n\n\n<li><a href=\"https:\/\/ai.meta.com\/blog\/llama-2-updates-connect-2023\/\" rel=\"noreferrer noopener\" target=\"_blank\">The Llama Ecosystem: Past, Present, and Future<\/a>. Llama 2, released by Meta, aims to broaden access to state-of-the-art AI technology. It brings value through research collaboration, enterprise insights, and leveraging emerging AI advancements.<\/li>\n\n\n\n<li><a href=\"https:\/\/huggingface.co\/blog\/Llama2-for-non-engineers\" rel=\"noreferrer noopener\" target=\"_blank\">Non-engineers guide: Train a LLaMA 2 chatbot<\/a>. Hugging Face offers a no-code solution for AI practitioners to build, train, and deploy chatbots. With tools like AutoTrain, ChatUI, and Spaces, even non-ML specialists can create advanced ML models, fine-tune LLMs, and easily interact with open-source LLMs. Spaces also simplifies the deployment of pre-configured ML applications and custom ML apps.<\/li>\n\n\n\n<li><a href=\"https:\/\/stratechery.com\/2023\/ai-hardware-and-virtual-reality\/\" rel=\"noreferrer noopener\" target=\"_blank\">AI, Hardware, and Virtual Reality<\/a>. This content explores the potential of AI, hardware, and virtual reality (VR). It discusses how AI challenges human limitations, enabling tasks that previously required human attention. It also highlights Meta\u2019s focus on AI integration in smart glasses to enhance user experience.<\/li>\n\n\n\n<li><a href=\"https:\/\/hbsp.harvard.edu\/inspiring-minds\/student-use-cases-for-ai\" rel=\"noreferrer noopener\" target=\"_blank\">Student Use Cases for AI<\/a>. Generative AI tools and Large Language Models (LLMs) have the potential to bring significant changes to education. While they empower students and educators with advanced technology, they also present challenges like the need for user verification and potential biases.<\/li>\n<\/ul>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"73e4\">\ud83d\udd2c Interesting Papers and Repositories<\/h1>\n\n\n\n<ul class=\"wp-block-list\">\n<li><a href=\"https:\/\/arxiv.org\/abs\/2309.14717\" rel=\"noreferrer noopener\" target=\"_blank\">QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models<\/a>. QA-LoRA, a new method in quantization-aware training, outperforms QLoRA in terms of efficiency and accuracy. It effectively balances the trade-off between quantization and adaptation, leading to minimal accuracy loss. QA-LoRA is especially effective in low-bit quantization scenarios like INT2\/INT3, without the need for post-training quantization, and it can be applied to various model sizes and tasks.<\/li>\n\n\n\n<li><a href=\"https:\/\/arxiv.org\/abs\/2309.14322\" rel=\"noreferrer noopener\" target=\"_blank\">Small-scale proxies for large-scale Transformer training instabilities<\/a>. A study has discovered that instabilities in training large Transformer-based models can be detected in advance by analyzing activations and gradient norms. These instabilities, which occur in both smaller and larger models with higher learning rates, can be mitigated using strategies employed in large-scale settings.<\/li>\n\n\n\n<li><a href=\"https:\/\/arxiv.org\/abs\/2309.16235\" rel=\"noreferrer noopener\" target=\"_blank\">Language models in molecular discovery<\/a>. Language models are being used in chemistry to accelerate the process of molecule discovery and show potential in early-stage drug research. These models assist in de novo drug design, property prediction, and reaction chemistry, offering a faster and more effective approach to the field. Moreover, open-source software for language modeling is enabling scientists to easily access and advance scientific language modeling, facilitating quicker chemical discoveries.<\/li>\n\n\n\n<li><a href=\"https:\/\/github.com\/TabbyML\/tabby\" rel=\"noreferrer noopener\" target=\"_blank\">Tabby: Self-hosted AI coding assistant<\/a>. Tabby is a fast and efficient open-source AI coding assistant compatible with popular language models. It supports CPU and GPU for coding tasks and offers swift coding experiences.<\/li>\n\n\n\n<li><a href=\"https:\/\/arxiv.org\/abs\/2309.12288v2\" rel=\"noreferrer noopener\" target=\"_blank\">The Reversal Curse: LLMs trained on \u201cA is B\u201d fail to learn \u201cB is A\u201d<\/a>. Researchers have discovered a phenomenon called the \u201cReversal Curse\u201d that affects the generalization abilities of auto-regressive language models (LLMs). These models struggle to infer the reverse of a fact, hindering their ability to answer related questions accurately. Even larger models like GPT-3.5 and GPT-4 face challenges in addressing this issue, indicating a need for further advancements in language modeling.<\/li>\n\n\n\n<li><a href=\"https:\/\/arxiv.org\/abs\/2309.11523\" rel=\"noreferrer noopener\" target=\"_blank\">RMT: Retentive Networks Meet Vision Transformers<\/a>. The Retentive Network (RetNet) has gained attention in the NLP community and shows potential as a replacement for Transformers. The combination of RetNet and Transformer, known as RMT, achieves outstanding results in vision tasks, with high performance metrics and surpassing existing vision backbones in downstream tasks.<\/li>\n\n\n\n<li><a href=\"https:\/\/arxiv.org\/abs\/2309.16534\" rel=\"noreferrer noopener\" target=\"_blank\">MotionLM: Multi-Agent Motion Forecasting as Language Modeling<\/a>. MotionLM is a new model that uses language processing to accurately predict the movements of multiple cars on the road. It outperforms other models by generating joint distributions and ranking the future interactions of agents in a single decoding process. MotionLM has demonstrated its effectiveness by securing the top spot on the Waymo Open Motion Dataset challenge leaderboard.<\/li>\n\n\n\n<li><a href=\"https:\/\/arxiv.org\/abs\/2309.16058\" rel=\"noreferrer noopener\" target=\"_blank\">AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model<\/a>. Introducing AnyMAL, a multimodal model designed to process diverse input signals including text, image, video, audio, and IMU motion sensor data. With its powerful text-based reasoning capabilities and a pre-trained aligner module, AnyMAL effectively understands and processes varied inputs. It is fine-tuned with a multimodal instruction set, expanding its capabilities beyond traditional question-answer scenarios.<\/li>\n<\/ul>\n\n\n\n<p id=\"001f\">Thank you for reading! If you want to learn more about NLP, remember to follow&nbsp;<a href=\"https:\/\/www.nlplanet.org\/\" rel=\"noreferrer noopener\" target=\"_blank\">NLPlanet<\/a>. You can find us on&nbsp;<a href=\"https:\/\/www.linkedin.com\/company\/nlplanet\" rel=\"noreferrer noopener\" target=\"_blank\">LinkedIn<\/a>,&nbsp;<a href=\"https:\/\/twitter.com\/nlplanet_\" rel=\"noreferrer noopener\" target=\"_blank\">Twitter<\/a>,&nbsp;<a href=\"https:\/\/medium.com\/nlplanet\">Medium<\/a>, and our&nbsp;<a href=\"https:\/\/discord.gg\/zfC862H2dJ\" rel=\"noreferrer noopener\" target=\"_blank\">Discord server<\/a>!<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Voice and image capabilities into ChatGPT, Amazon invests 4B$ in Anthropic, and Mistral LLM Here are your weekly articles, guides, and news about NLP and AI chosen for you by\u00a0NLPlanet!&#8230; <a class=\"read-more-link\" href=\"https:\/\/tbekk.com\/devstream\/2023\/10\/05\/weekly-ai-and-nlp-news-october-2nd-2023\/\">Read more &raquo;<\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[181,51,115,201,132,9],"tags":[255,302,315,316],"class_list":["post-860","post","type-post","status-publish","format-standard","hentry","category-ai-2","category-article","category-data-science","category-nlp","category-nn","category-news","tag-language-models","tag-news","tag-qa-lora","tag-vision-transformers"],"_links":{"self":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/860","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/comments?post=860"}],"version-history":[{"count":3,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/860\/revisions"}],"predecessor-version":[{"id":864,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/860\/revisions\/864"}],"wp:attachment":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/media?parent=860"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/categories?post=860"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/tags?post=860"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}