Apple has its eyes on a technology that could seriously change how artificial intelligence works on the iPhone.
And we are not talking about some minor tweak here.
The company held meetings with the startup PrismML to explore how it could use their technology to run larger AI models directly on devices, without relying on external servers.
The information came from The Information and has already sparked a lot of discussion in the tech world.
What really stands out is what PrismML managed to pull off in practice:
- Ran the Qwen 3.6 model, from Alibaba, directly on an iPhone 17 Pro
- The model has 27 billion parameters, all active at the same time
- This goes beyond what Apple’s current on-device model can do
If that sounds too technical right now, don’t worry. We are going to break it all down in a simple way just below. 👇
What is PrismML and why does it matter so much here?
PrismML is a startup that develops optimization technology to run heavy artificial intelligence models directly on hardware with limited resources, like smartphones. Instead of relying on cloud servers to process data, their approach is to make all of that work happen right there in your pocket. It sounds simple to explain, but in practice it is a massive technical challenge because large AI models consume a lot of memory, processing power, and energy — exactly the things a phone has in very limited supply.
The demo that caught Apple’s attention was running Qwen 3.6, an open-source language model created by Alibaba, entirely inside an iPhone 17 Pro. The impressive detail is not just the fact that the model ran on the device, but rather its size: 27 billion parameters, all active simultaneously. To give you an idea, parameters are like the internal connections and adjustments that allow an AI model to think and respond. The more parameters, the more capable and sophisticated the model tends to be. And running all of that in a dense format, without turning off parts of the model to save resources, is something very few technologies can do on mobile devices today.
This achievement puts PrismML in a very strategic position. Major tech companies like Apple are constantly looking for ways to improve the artificial intelligence experience in their products without compromising user privacy. When a small startup shows up with a solution that seems to solve exactly that problem, the conversations happen fast. And apparently, that is what happened here. 🤝
Understanding the difference between dense and sparse architecture
This is the most interesting point of this whole story, and it is worth explaining carefully. The model that Apple currently uses on-device, called AFM 3 Core Advanced, has 20 billion parameters. But it works with a sparse architecture, a pretty clever way to save resources. In this type of architecture, only a small portion of the model is active at any given moment, somewhere between 1 billion and 4 billion parameters at a time. It is like having a huge team but only assigning a few specialists to each specific task while keeping the rest on standby.
What PrismML managed to do with Qwen 3.6 is different and much more ambitious. All 27 billion parameters stay active at the same time in a dense architecture. That means the model’s full capacity is available at all times, with no shortcuts or cutbacks to save processing power. This would normally require a lot more memory and computing power — things a smartphone has in very limited amounts. Making that fit and actually work inside an iPhone 17 Pro is exactly the kind of technical breakthrough that makes Apple engineers light up. 💡
Why does Apple want larger models inside the iPhone?
Apple launched Apple Intelligence with a clear promise: to bring artificial intelligence in an integrated, private, and efficient way to its devices. The idea was exactly that — process as much as possible directly on the device without sending your data to third-party servers. This is a huge advantage in terms of privacy, and the company made sure to communicate it with a lot of emphasis. The problem is that keeping that promise while delivering an increasingly capable AI is a tough equation to solve. Larger models mean smarter responses, more context, and more accuracy. But they also demand more from the hardware.
The model that Apple currently runs locally on the iPhone, the AFM 3 Core Advanced, is exactly what powers some of the coolest features in iOS 27. We are talking about the more expressive Siri AI voices and the improved system-wide dictation, available on the iPhone 17 Pro and iPhone Air models. These features already show how much a good on-device AI makes a difference. Now imagine taking that to an even higher level with a model that has all of its parameters working simultaneously. The experience would become much richer and more sophisticated.
That is why Apple is exploring technologies like PrismML’s. The company wants to keep delivering AI on the device but wants to take a big leap in the quality of that experience. If the startup’s technology allows much more powerful models to run without needing a cloud connection, that solves two problems at once: it keeps user privacy intact and also reduces the costs Apple has with its server infrastructure. It is a combination that makes perfect sense within the company’s product strategy. 📱
The role of Private Cloud Compute in this story
Today, when an Apple Intelligence feature is too heavy to run on the device, it gets sent to Private Cloud Compute servers, Apple’s private cloud infrastructure. This system was designed to maintain privacy even when processing happens off-device, which is already a major step forward compared to the competition. Even so, keeping those servers running is expensive, and every request that leaves the phone generates costs and depends on an internet connection.
If Apple can run larger models directly on the iPhone, more artificial intelligence features could work locally, reducing the dependence on Private Cloud Compute. That means lower operational costs for Apple and an extra layer of privacy for the user, since the data would not even need to leave the device. It is a scenario where everybody wins, and that is exactly why this possibility is being taken so seriously by the company.
What changes for everyday iPhone users?
When we talk about larger artificial intelligence models running directly on the iPhone, the practical impact goes well beyond technical numbers. Think about simple situations: asking Siri to summarize a long conversation in WhatsApp, generating a more polished email draft, or having a writing assistant that actually understands the context of what you are typing. Today, many of these tasks are limited precisely by the capabilities of the model running locally. With a jump to something in the range of 27 billion active parameters, the quality of responses changes in a very concrete way — becoming smoother, more accurate, and more useful.
Another important point is latency, which is the time it takes for the model to respond after you ask a question or give a command. When artificial intelligence needs to reach an external server to process your request, there is a wait time that depends on your internet connection. In areas with poor signal or no Wi-Fi, this can slow things down or even prevent features from working at all. With everything processing inside the iPhone itself, the response would be nearly instant, no matter where you are. This is especially relevant when you are on the go, traveling, or in situations where connectivity is unreliable.
And speaking of privacy, this might be the most important benefit for the average user, even if it is the least visible one. When data never leaves the device, no company, server, or external service has access to what you are typing, asking, or processing. For people who use AI for sensitive things — like organizing personal finances, writing professional documents, or handling health information — knowing that everything stays on the device makes a huge difference. Apple understood this before many competitors, and now it seems to be doubling down in that direction with help from PrismML’s technology. 🔒
Is this partnership already confirmed?
Not yet. What The Information reported is that Apple held meetings with PrismML to explore the use of the startup’s technology. Meetings like this are common in the tech world, especially when a major company wants to understand what is available on the market before making a decision. Apple could end up acquiring the startup, licensing the technology, or simply using these conversations to better understand what is possible and develop a solution in-house. None of these paths are ruled out at this point.
What already says a lot, however, is the simple fact that these meetings took place. Apple is known for being extremely selective and discreet with its strategic moves. When the company sits down to talk with a startup, it is because something there genuinely caught their attention. And the demo of Qwen 3.6 running on the iPhone 17 Pro with all 27 billion parameters active simultaneously is, without question, the kind of thing that gets any AI engineer at the company very interested. PrismML’s optimization technology solves a problem that Apple is clearly trying to overcome, and that alone is a strong signal.
The artificial intelligence market is moving at a breakneck pace, and the companies that manage to run larger models efficiently inside their own devices will have a real competitive advantage in the coming years. Apple has already shown it takes this seriously, and moves like this reinforce that the company is looking beyond what the iPhone does today, thinking about how it will work tomorrow. It is going to be really interesting to follow the next chapters of this story. 👀
