Guide to Installing and Running LLMs Locally on Android Phones

Guide to Installing and Running LLMs Locally on Android Phones

In the realm of artificial intelligence, where apps like LM Studio and GPT4All reign supreme on desktops, the landscape on Android devices is a barren one. However, fear not, for MLC LLM has bestowed upon us a gift in the form of MLC Chat – an Android app that allows the local deployment of LLM models. Embark on this journey with me as we delve into the world of possibility.

Behold, dear reader, a caveat to heed: MLC Chat, in its current state, does not fully utilize the on-device NPU in all Snapdragon devices, resulting in sluggish token generation. Alas, the burden falls upon the CPU for inference. Yet, fret not, for certain devices, such as the Samsung Galaxy S23 Ultra (powered by Snapdragon 8 Gen 2), have been optimized for the MLC Chat app, promising a smoother experience.

As we chart our course into the unknown, I implore you to take the first step – download the MLC Chat app for your Android phone. A mere click, a download of the APK file (148MB), and installation await you. Embrace the AI models that beckon, from Llama 3 to Gemma, Phi-2 to Mistral, and beyond.

With bated breath, launch the MLC Chat app and witness the array of AI models at your disposal. Should you seek the cutting-edge Llama 3 8B model or the nimble Phi-2, the choice is yours. As for me, I chose the Microsoft’s Phi-2 model – compact and agile, a companion on this journey.

Once the die is cast, tap the chat button beside the chosen model. Behold, as the symphony of words unravels before your eyes – a conversation with the AI model, all within the confines of your Android device, devoid of the shackles of an internet connection.

In the crucible of testing, the Phi-2 model graced my phone with its presence, though marred by occasional hallucinations. Alas, Gemma’s spirit remained elusive, and the behemoth that is Llama 3 8B faltered in its stride.

A revelation, my friends, awaits those with Snapdragon-powered devices. Witness the Snapdragon 855+ SoC in my OnePlus 7T, weaving its magic at a pace of three tokens per second whilst running Phi-2. A testament to the power that lies within.

And thus, we conclude our tale of local LLM model deployment on Android devices. Though the path may be wrought with challenges, the promise of AI models on Android phones is a beacon of hope. The era of the CPU may soon give way to a symphony of NPU, GPU, and CPU harmonizing in Qualcomm’s AI Stack implementation.

For our comrades on the Apple side, the MLX framework stands as a bastion of quick local inferencing on iPhones, generating eight tokens per second with aplomb. A portent of things to come for Android devices, as they too may soon harness the on-device NPU, ushering in an era of unparalleled performance.

As we bid adieu, dear reader, remember – should you wish to commune with your documents through a local AI model, our dedicated article shall guide you. And should you encounter tribulations along the way, fear not, for the comment section below beckons, eager to lend a helping hand.

Support our work ❤️

If you enjoyed this article, consider leaving a tip to help us keep publishing great content.

Secure payment on PayPal
See also:  Volkswagen-Licensed n+ E-Bikes Feature Smart Glasses
Moyens I/O Staff is a team of expert writers passionate about technology, innovation, and digital trends. With strong expertise in AI, mobile apps, gaming, and digital culture, we produce accurate, verified, and valuable content. Our mission: to provide reliable and clear information to help you navigate the ever-evolving digital world. Discover what our readers say on Trustpilot.