In the epoch of OpenAI’s ChatGPT, chatbots have taken center stage as the quintessence of Artificial Intelligence (AI) in our contemporary world. It seems relishing dialogues with an AI chatbot is the solitary mode of engaging with AI models and cerebral systems. Yet, I contend that confining your aspirations of interacting with an intelligent system within the confines of a text chatbox is a fallacy.
In this milieu, Microsoft has plunged into the whirlwind of infusing AI chatbots into a myriad of its products. Foremost among them is the integration of Windows Copilot, an AI chatbot fueled by OpenAI’s models, into Windows 11 with extravagant fanfare. Noteworthy is the fact that Microsoft has supplanted Cortana with Windows Copilot on Windows 11, and similarly embedded Windows Copilot into Windows 10, ousting Cortana in the process.
Undoubtedly, Microsoft envisions AI chatbots as the vanguard of the future. But is it genuinely the epitome of intelligent computing propelled by AI, or is Microsoft merely pandering to the AI frenzy, incorporating AI chatbots to demonstrate to investors that it is a contender in the arena? Regardless of the answer, the current iteration of AI-driven chatbots exhibits a restricted application, rendering it arduous to glean substantial assistance from the chatbot, particularly at the operating system level.
Windows Copilot: A Deterioration from Cortana?
Microsoft opted to phase out Cortana—a product with a long-standing nine-year presence—in favor of Windows Copilot. But does Windows Copilot serve as a fitting substitute, particularly while still in its nascent preview phase?
Nonetheless, let us delve into a point-by-point comparison. Firstly, Cortana predominantly functioned as a voice assistant, whereas Windows Copilot is a text-based AI chatbot, albeit it accommodates voice input as an option but not by default.
In essence, Windows Copilot is not tailored for a voice-centric user experience, which engenders a disjointed user interaction, unlike the personal touch offered by Cortana. The user-friendliness and intuitive nature of voice input prevail over text input in terms of UI approachability, rendering Windows Copilot deficient in the fundamental domain of user experience.
When it comes to features, Cortana had evolved into a robust product with the capability of executing a plethora of system-level functions. It could set timers, alarms, reminders, compose emails, define terms, launch applications, and execute an array of additional tasks. Essentially, Cortana was deeply interwoven into the Windows OS fabric, possessing a profound understanding of the system.
Conversely, Copilot operates on general-purpose large language models (LLM) which are not optimized for executing localized functions on Windows. When requesting Windows Copilot to set a timer, it redirects to an online service for the task. It falls short on basic functions like setting alarms or playing music, merely launching the Spotify app. In essence, there seems to be an absence of any substantial AI prowess in Copilot.
Microsoft appears to be hastily hopping aboard the AI bandwagon, reminiscent of its regret over missing the smartphone race, aiming to steer clear of repeating the same oversight.
Of course, Windows Copilot is still in its nascent phase, and these functionalities are likely to be introduced in the future (some are currently being tested in Insider builds). But what warranted the swift substitution of a nearly decade-old product like Cortana with a rudimentary AI chatbot?
It seems to me as though Microsoft has haphazardly integrated Copilot and deemed it sufficient, at least for the time being. The tech giant has not endeavored to bridge the gap between Copilot and Cortana in terms of feature parity before replacing the longstanding product.
It’s especially disheartening because Microsoft is introducing a Copilot key on the Windows keyboard—something Microsoft describes as a “significant change to the Windows PC keyboard in nearly three decades”—yet minimal thought has gone into it.
Where is the AI Wizardry in Windows Copilot?
Coming to the functionalities of Windows Copilot, one can pose inquiries on any topic and receive immediate responses. Additionally, users can transition to the Creative mode to engage with the potent GPT-4 model.
Windows Copilot can succinctly summarize a webpage, unearth pivotal insights, devise itineraries, and more. Microsoft has integrated a screenshot tool into Copilot utilizing the GPT-4V model for visual analysis. This tool can be utilized for OCR tasks or obtaining information about images.
In terms of Windows-specific features, users can address issues like audio problems, following which Copilot guides them through the audio troubleshooter. It aids in troubleshooting other Windows dilemmas as well. Furthermore, users can toggle dark mode, capture screenshots, and arrange windows through Copilot.
While these functionalities are commendable for the preview version of Windows Copilot, most of them also function in Edge Copilot, except for the Windows-specific attributes. Moreover, Windows Copilot is unable to access webpages from browsers other than Chrome. Since it operates on Edge’s engine, it lacks access to content from alternate windows, whether browsers, Notepad, or Office applications.
This presents a glaring void in the implementation of Windows Copilot. Rather than being developed using the WinUI 3 framework for delivering a native experience, Copilot functions as an auxiliary of the Edge browser. Consequently, the deep integration of Windows Copilot into pivotal elements of the OS is notably absent.
For instance, users cannot right-click on a file in Windows Explorer and request Copilot to elucidate it, convert the file format, or execute any desired action. It would be remarkable if users could submit an Excel file to Copilot from the context menu for real-time data analysis. Presently, apart from images, there is no viable method to interact with files utilizing Windows Copilot on Windows 11.
Windows Copilot: A Case of Overpromise and Underdelivery
Lately, Microsoft has been adept at unveiling and marketing novel features, yet when it comes to utilizing the promised features, they are conspicuously absent. Three months ago, when Windows Copilot was announced, it pledged several new functionalities, which are either currently unavailable or do not operate as advertised.
For instance, when users ask Copilot to organize their windows, it solicits permission and only arranges one window, leaving users to manage the remaining actions. Similarly, it fails to play music tailored to specific moods when prompted. Instead, Copilot merely generates links from platforms like YouTube, diverging from the expectations associated with an intelligent AI-powered Copilot.
Furthermore, the eagerly awaited contextual menu for Copilot has not been introduced as yet. Features like Rewrite, Explain, and Summarize are absent for any active window. The absence of Draft with Copilot is conspicuous even three months post-release. Not to mention, the removal of image backgrounds and the addition of Extension support are yet to materialize.
The touted and hyped-up features seem to be conspicuously absent. Microsoft’s track record bespeaks of overpromising and underdelivering with numerous of its products.
What Could Be the Vision for Windows Copilot?
Examining the potential trajectory of Windows Copilot, it is intriguing to contemplate what the open-source community is spearheading. Notably, the Open Interpreter tool has emerged as a captivating utility facilitating interaction with local files, conversion to diverse formats, handling various file formats, generating charts, and engendering various actions on Windows.
Just recently, the latest iteration of Open Interpreter (0.2.0) surfaced with a spellbinding OS mode. Users can manipulate their computers with uncomplicated natural language prompts. Open Interpreter leverages vision models such as GPT-4V to comprehend the GUI environment and execute actions on users’ computers.
For instance, users can instruct it to activate dark mode, prompting it to launch the appropriate Settings page and toggle the setting utilizing the Vision model.
These rudimentary examples underscore the capabilities of vision models, while Windows Copilot remains stagnant in presenting textual information through a chatbox, far from embodying true intelligence.
A truly astute Copilot should have the capacity to dispatch emails, fine-tune Windows settings, interact with the OS at the system level, and undertake an array of other functionalities. The potential utility is boundless, offering considerable benefits in enhancing accessibility on Windows 11 24H2.
Undoubtedly, invoking the GPT-4V API would entail a significant outlay for Microsoft. Yet, devising a compact vision model explicitly for Windows, akin to CogVLM, could abate latency, facilitating ubiquitous execution, even when PC connectivity is disrupted.
With forthcoming Intel and Snapdragon X Elite chipsets incorporating dedicated NPUs, supporting the operation of smaller models on-device becomes feasible. Whether Microsoft opts for running its in-house developed visual model in the cloud or transitioning to an on-device framework, the costs would be substantially lower.
Introducing r1. Watch the keynote.
Order now: https://t.co/R3sOtVWoJ5 #CES2024 pic.twitter.com/niUmjFvKvE— rabbit inc. (@rabbit_hmi) January 9, 2024
Illustrating an alternative example, the recent demonstration of Rabbit R1—an AI-centric hardware apparatus—showcases its capability to execute diverse tasks. Propelled by what they describe as a Large Action Model (LAM), Rabbit R1 adeptly performs tasks like ordering pizza, dispatching emails, and booking flights with a modicum of voice input.
Microsoft should contemplate deploying an analogous LAM designed for executing actions, transcending mere conversational interactions with a chatbot.
If a budding startup such as Rabbit succeeds in this endeavor, a behemoth like Microsoft, laden with abundant resources, should be more than capable. Thus far, Microsoft has charted its course by crafting the Phi-2 model, a diminutive LLM designed solely for research purposes. If Microsoft aspires to offer an AI-integrated PC experience in 2024, it must fashion Windows-specific vision models to execute actions locally devoid of latency. Microsoft needs to concoct a version of the GPT-4V API tailored for local execution on Windows, propelling the functionality to new heights.
Windows Copilot Demands a Fresh Approach
To culminate, Windows Copilot, in its present chatbot configuration, encompasses a narrow scope of potential applications, an arena already saturated with a myriad of browser extensions. Microsoft necessitates a novel approach to orchestrate the conception of AI-powered PCs.
Microsoft’s fierce rival, Apple, is renowned for meticulously crafting products, waiting until they are perfected before unveiling them to the public. Conversely, Microsoft’s strategy veers in the opposite direction, releasing products prematurely lacking essential and substantive features at launch.
Symbolic of Microsoft’s approach towards AI without due deliberation, branding Edge browser as an “AI browser” solely by embedding a chatbot underscores a hasty tactic. The company endeavors to infuse AI attributes into Notepad, continues ameliorating AI-powered features within Office applications, MS Paint, Snipping Tool, and various proprietary apps.
Microsoft must transcend the fixation on integrating a chatbot and commence afresh.
While these in-app AI features may cater to a subset of users, the aspiration to fashion Windows as an inherently intelligent OS powered by AI mandates Microsoft to transcend the fixation on enveloping a chatbot within its ecosystem and chart innovative pathways that pave the way for novel and transformative AI-driven experiences.
Support our work ❤️
If you enjoyed this article, consider leaving a tip to help us keep publishing great content.


























