Category: Text to Music

AI race is heating up: Announcements by Google/DeepMind, Meta, Microsoft/OpenAI, Amazon/Anthropic

September 22, 2023 / admin / 0 Comments

After weeks of “less exciting” news in the AI space since the release of Llama 2 by Meta on July 18, 2023, there were a bunch of announcements in the last few days by major players in the AI space:

Google/DeepMind: Bard extensions and multimodal LLM Gemini
OpenAI: DALL-E3 and GPT-Vision in ChatGPT, Gobi
Microsoft: Windows Copilot with DALL-E3 access
Amazon: Generative AI in Alexa, $4B investment in Anthropic
Meta: Meta AI, Ray-Ban, Emu, AI studio

Here are some links to the news of the last weeks:

Sep 28, 2023, Amazon: Securely customize CodeWhisperer
Sep 27, 2023, Meta: Meta AI assistant, Ray-Ban smart glasses, Emu, AI studio
Sep 25, 2023, ChatGPT can now see, hear, and speak
Sep 25, 2023, Amazon invests $4B in Anthropic, Claude in Bedrock
Sep 25, 2023, Spotify clones voices and translates them
Sep 21, 2023, Announcing Microsoft Copilot
Sep 20, 2023, Amazon brings generative AI to Alexa
Sep 20, 2023, OpenAI Announces DALL·E 3 in Research Preview
Sep 20, 2023, GitHub Copilot Chat beta now available for all individuals
Sep 19, 2023, Google Bard September update: App extensions
Sep 19, 2023, OpenAI’s multimodal LLM GPT-Vision to beat Google Gemini
Sep 16, 2023, DeepMind: LLMs can optimize their own prompts
Sep 15, 2023, Google nears release of AI software Gemini
Sep 15, 2023, Google Gemini: What We Know So Far
Sep 13, 2023, Stable Audio by Stability AI for music & sound generation
Sep 07, 2023, Anthropic introduces Claude Pro
Sep 06, 2023, Falcon 180B
Aug 31, 2023, Baidu launches Ernie chatbot
Aug 29, 2023, Duet AI for Google Workspace now generally available
Aug 28, 2023, Meta plans to take on GPT-4 with a rumored Llama 3
Aug 28, 2023, Introducing ChatGPT Enterprise
Aug 27, 2023, Google Gemini Smashes GPT-4 By 5X
Aug 24, 2023, Introducing Code Llama
Aug 22, 2023, GPT-3.5 Turbo fine-tuning and API updates
Aug 22, 2023, ElevenLabs releases Eleven Multilingual v2
Aug 21, 2023, MidJourney Adds Inpainting Feature
Aug 16, 2023, Adobe Express with AI Firefly app is released worldwide
Aug 10, 2023, ChatGPT expands its ‘custom instructions’ feature
Aug 08, 2023, Announcing StableCode — Stability AI
Aug 05, 2023, Tim Cook says Apple is building AI into ‘every product’
Aug 03, 2023, Every single Amazon team is working on generative AI
Aug 02, 2023, AudioCraft by Meta
Jul 31, 2023, ChatGPT for Android in all countries

Google revealed PaLM 2

May 16, 2023 / admin / 0 Comments

Google revealed at Google I/O on May 10, 2023, PaLM 2 (API, paper), its latest AI language model that powers 25 Google products, including Search, Gmail, Docs, Assistant, Translate, and Photos.

PaLM 2 has 4 models that differ in size: Gecko, Otter, Bison, and Unicorn. Gecko is so lightweight that it can work on mobile devices.
PaLM 2 can be finetuned on domain-specific knowledge (Sec-PaLM with security knowledge, Med-PaLM 2 with medical knowledge)
Bard now works with PaLM 2; with extensions, Bard can call tools like Sheets, Colab for coding, Lenses, Maps, Adobe Firefly to create images, etc.; Bard is multimodal and can understand images
PaLM 2 is also powering Duet AI for Google Cloud, a generative AI collaborator designed to help users learn, build and operate faster
PaLM 2 is released in 180+ regions and countries, however, e.g. not yet in Canada, and in the EU
The next model, Gemini, is already in training.
Google also announced the availability of MusicLM, a text-to-music generative model.

OpenAI reacted to this announcement on May 12 by announcing that Browsing & Plugins are rolled out over the subsequent week for all Plus users. As of May 17, I can confirm that both features are now operational for me.

3rd-Level of Generative AI

April 11, 2023 / admin / 0 Comments

Defining

1st-level generative AI as applications that are directly based on X-to-Y models (foundation models that build a kind of operating system for downstream tasks) where X and Y can be text/code, image, segmented image, thermal image, speech/sound/music/song, avatar, depth, 3D, video, 4D (3D video, NeRF), IMU (Inertial Measurement Unit), amino acid sequences (AAS), 3D-protein structure, sentiment, emotions, gestures, etc., e.g.

_{X = text, Y = text: LLM-based chatbots like ChatGPT (from OpenAI based on LLMs GPT-3.5 [4K context] or GPT-4 [8K/32K context]), Bing Chat (GPT-4), Bard (from Google, based on PaLM 2), Claude (from Anthropic [100K context]), Llama2 (from Meta), Falcon 180B (from Technology Innovation Institute), Alpaca, Vicuna, OpenAssistant, HuggingChat (all based on LLaMA [GitHub] from Meta), OpenChatKit (based on EleutherAI’s GPT-NeoX-20B), CarperAI, Guanaco, My AI (from Snapchat), Tingwu (from Alibaba based on Tongyi Qianwen), (other LLMs: MPT-7B and MPT-30B from Mosaic [65K context, commercially usable], Orca, Open-LLama-13b), or coding assistants (like GitHub Copilot / OpenAI Codex, AlphaCode from DeepMind, CodeWhisperer from Amazon, Ghostwriter from Replit, CodiumAI, Tabnine, Cursor, Cody (from Sourcegraph), StarCoder from Big Code Project led by Hugging Face, CodeT5+ from Salesforce, Gorilla, StableCode from Stability.AI, Code Llama from Meta), or writing assistants (like Jasper, Copy.AI), etc.}
_{X = text, Y = image: Dall-E (from OpenAI), Midjourney, Stable Diffusion (from Stability.AI), Adobe Firefly, DeepFloyd-IF (from Deep Floyd, [GitHub, HuggingFace]), Imagen and Parti (from Google), Perfusion (from NVIDIA)}
_{X = text, Y = 360° image: Skybox AI (from Blockade Labs)}
_{X = text, Y = 3D avatar: Tafi}
_{X = text, Y = avatar lip sync: Ex-Human, D-ID, Synthesia, Colossyan, Hour Once, Movio, YEPIC-AI, Elai.io}
_{X = speech + face video, Y = synched audio-visual: Lalamu}
_{X = text, Y = video: Gen-2 (from Runway Research), Imagen-Video (from Google), Make-A-Video (from Meta), or from NVIDIA}
_{X = text, Y = video game: Muse & Sentis (from Unity)}
_{X = image, Y = text: GPT-4 (from OpenAI), LLaVA}
_{X = image, Y = segmented image: Segment Anything Model (SAM by Meta)}
_{X = speech, Y = text: STT (speech-to-text engines) like Whisper (from OpenAI), MMS [GitHub] (from Meta), Conformer-2 (from AssemblyAI)}
_{X = text, Y = speech: TTS (text-to-speech engines) like VALL-E (from Microsoft), Voicebox (from Meta), SoundStorm (from Google), ElevenLabs, Bark, Coqui}
_{X = text, Y = music: MusicLM (from Google), RIFFUSION, AudioCraft (MusicGen, AudioGen, EnCodec from Meta), Stable Audio (from Stability.ai)}
_{X = text, Y = song: Voicemod}
_{X = text, Y = 3D: DreamFusion (from Google)}
_{X = text, Y = 4D : MAV3D (from Meta)}
_{X = image, Y = 3D : CSM}
_{X = image, Y = audio: ImageBind [1] (from Meta, on GitHub)}
_{X = audio, Y = image: ImageBind [2] (from Meta)}
_{X = music, Y = image: MusicToImage}
_{X = text, Y = image & audio: ImageBind [3] (from Meta)}
_{X = audio & image, Y = image: ImageBind [4] (from Meta)}
_{X = IMU, Y = video: ImageBind (from Meta)}
_{X = AAS, Y = 3D-protein: AlphaFold (from Google), RoseTTAFold (from Baker Lab), ESMFold (from Meta)}
_{X = 3D-protein, Y = AAS: ProteinMPNN (from Baker Lab)}
_{X = 3D structure, Y = AAS: RFdiffusion (from Baker Lab)}

and 2nd-level generative AI that builds some kind of middleware and allows to implement agents by simplifying the combination of LLM-based 1st-level generative AI with other tools via actions (like web search, semantic search [based on embeddings and vector databases like Pinecone, Chroma, Milvus, Faiss], source code generation [REPL], calls to math tools like Wolfram Alpha, etc.), by using special prompting techniques (like templates, Chain-of-Thought [COT], Self-Consistency, Self-Ask, Tree Of Thoughts, ReAct [Reason + Act], Graph of Thoughts) within action chains, e.g.

_{ChatGPT Plugins (for simple chains)}
_{LangChain + LlamaIndex (for simple or complex chains)}
_ToolFormer

we currently (April/May/June 2023) see a 3rd-level of generative AI that implements agents that can solve complex tasks by the interaction of different LLMs in complex chains, e.g.

_BabyAGI
_Auto-GPT
_{Llama Lab (llama_agi, auto_llama)}
_{Camel, Camel-AutoGPT}
_{JARVIS (from Microsoft)}
_{Generative Agents}
_{ACT-1 (from Adept)}
_Voyager
_SuperAGI
_{GPT Engineer}
_Parsel
_MetaGPT

However, older publications like Cicero may also fall into this category of complex applications. Typically, these agent implementations are (currently) not built on top of the 2nd-level generative AI frameworks. But this is going to change.

Other, simpler applications that just allow semantic search over private documents with a locally hosted LLM and embedding generation, such as e.g. PrivateGPT which is based on LangChain and Llama (functionality similar to OpenAI’s ChatGPT-Retrieval plugin), may also be of interest in this context. And also applications that concentrate on the code generation ability of LLMs like GPT-Code-UI and OpenInterpreter, both open-source implementations of OpenAI’s ChatGPT Code Interpreter/AdvancedDataAnalysis (similar to Bard’s implicit code execution; an alternative to Code Interpreter is plugin Noteable), or smol-ai developer (that generates the complete source code from a markup description) should be noticed.
There is a nice overview of LLM Powered Autonomous Agents on GitHub.

The next level may then be governed by embodied LLMs and agents (like PaLM-E with E for Embodied).

Google publishes MusicLM, a text to music language model

January 27, 2023 / admin / 0 Comments

Google Research published an impressive language model that can turn a text description into high-quality music [webpage, paper]. The source code is unfortunately not publicly available.

RIFFUSION: Stable Diffusion for Real-Time Music Generation

December 16, 2022 / admin / 0 Comments

By using the stable diffusion model v1.5 without any modifications, just fine-tuned on images of spectrograms paired with text, the software RIFFUSION (RIFF + diffusion) generates incredibly interesting music from text input. By interpolating in latent space it is possible to transition from one text prompt to the next. You can try out the model here.

The authors provide source code on GitHub for an interactive web app and an inference server. A model checkpoint is available on Hugging Face.

There is a nice video about RIFFUSION by Alan Thompson on youtube.

Even more shocking than using diffusion on spectrograms and getting great results may be a paper by Google Research published on Dec 15, 2022. They use text as an image and train their model with contrastive loss alone, thus calling their model CLIP-Pixels Only (CLIPPO). It’s a joint model for processing images and text with a single ViT (Vison Transformer) approach and astonishing performance.

Category: Text to Music

AI race is heating up: Announcements by Google/DeepMind, Meta, Microsoft/OpenAI, Amazon/Anthropic

Google revealed PaLM 2

3rd-Level of Generative AI

Google publishes MusicLM, a text to music language model

RIFFUSION: Stable Diffusion for Real-Time Music Generation

Recent Posts

Recent Comments

Archives

Categories