In the year of our digital domain 2022, when OpenAI unfurled the now-legendary ChatGPT, the realm of technology beheld a grand spectacle of innovation like no other. A marvel to behold, this beauteous creation heralded a new era of advancement that knows no bounds in its ceaseless expansion. Behold, for AI Chatbots have emerged from the chrysalis of giants such as Google, Microsoft, Meta, Anthropic, and other titans of industry. Powered by the mystical energies of LLMs – those enigmatic Large Language Models – these wondrous creations have taken flight into the boundless expanse of human language. But what manner of creature is this Large Language Model, and how does it toil in the vineyards of knowledge? Let us embark upon a quest of discovery as we delve into the mysteries of LLMs, guided by the ancient magics of our forebears.
In the arcane dialect of the technomancer, an LLM, or Large Language Model, emerges as a sublime creation of Artificial Intelligence, forged in the fiery crucible of a vast repository of texts. Designed to embody the very essence of human language, this entity of probability moves in the shadows of the ancient algorithms of deep learning. A wielder of arcane power, an LLM conjures forth essays, poems, articles, and letters; it breathes life into lines of code, transmutes texts from one tongue to another, and weaves intricate tapestries of summation. Verily I say unto thee, the larger the crucible of knowledge in which an LLM is forged, the more potent its incantations of Natural Language Processing become. Within the hallowed halls of AI research, sages doth proclaim that LLMs boasting 2 billion parameters or more are deemed as the “large” language models. But lo! Should thee wonder what manner of sorcery governs these parameters, let it be known that they are the very keystones upon which the model is hewn. With each passing moon, the denizens of technological lore unveil ever grander models, imbued with the incantations of larger parameters, granting them an undreamt of scope and power.
Behold! In days of yore, when OpenAI did unfurl the GPT-2 LLM upon the mortal realm in the year of 2019, it bore upon its brow the mantle of 1.5 billion parameters. Yet, in the fullness of time, did GPT-3 emerge from the depths of creation in 2020, possessing a staggering 175 billion parameters – a behemoth of over 116 times the size of its predecessor. And now, the very zenith of the craft, the magnum opus of LLMs, the state-of-the-art GPT-4, stands before us with a majestic countenance of 1.76 trillion parameters. Witness, dear reader, as the march of progress unfolds before us, each model greater and more magnificent than the last, blooming with the promise of advanced and ever more intricate capabilities.
In the luminous realm of revelation, LLMs embark upon a sacred quest of knowledge acquisition like unto the brightest stars in the heavens. Through the hallowed rites of pre-training, these entities of intellect learn to divine the very essence of language itself, seeking to prophesy the very words that shall come to pass. Within the crucible of a vast corpus of text, be it ancient lore from books, chronicles of news, or the endless scrolls of websites and Wikipedia, an LLM learns the secrets of grammar, syntax, the truths of the world, the art of reason, and the hidden patterns that bind the tapestry of words together in a grand symphony of wisdom. But lo, after the rites of pre-training are concluded, a model must venture forth into the refinement of fine-tuning, wherein it hones its skills upon specific datasources. Whether the incantation be for the mastery of code or the art of storytelling, the model must journey through the crucible of fine-tuning to attain its full glory.
In the annals of history, a mighty revolution did dawn upon the realm of LLMs in the year of 2017, when the wise sages of the Google Brain team did bring forth a seminal scroll known as “Attention is All You Need” by the esteemed scribes Vaswani et al. Within the sacred confines of this scroll lay the breath of life for a new era, the era of the Transformer architecture. A marvel to behold, this wondrous creation spawned from the very essence of self-attention, enabling a model to gaze upon all words within a sentence in parallel, grasping the very essence of their context and the threads that bind them. Verily, it unlocked the ancient magic of parallelism, heralding a new age of efficiency in the arts of training. As the fates did decree, Google unveiled the mighty BERT, the first transformer-based LLM, in the year of 2018. And lo, OpenAI did not tarry, for they too joined the chorus with their own GPT-1 model, embracing the transformative power of the transformer architecture.
Support our work ❤️
If you enjoyed this article, consider leaving a tip to help us keep publishing great content.


























