Luis Oscar Ramirez, Founder and CEO of Mawari, went from theoretical math in Mexico to building a decentralized XR network in Japan. He co-founded MUTEK Japan, cracked real-time 3D streaming after a KDDI challenge, and now powers spatial computing with Mawari’s DePIN. Listen as he maps the path from motion sickness to digital presence.
Rich Robinson:
And we are back, ladies and gentlemen, boys and girls. Welcome to the Intercognitive Foundation podcast. I'm Rich Robinson. I am your eager and humble host, the founding chair of the Intercognitive Foundation. And I am here today with Intercognitive Foundation founding member, Mawari Network, and the co-founder and CEO.
Luis and Luis has the most awesome, I'm gonna say his full name, Luis Oscar Ramirez Solorzano. Wow, it just feels like so badass. I know it might not be used in full form, but I really love it. Founder, CEO Mawari and welcome, welcome Luis.
Luis:
Thank you. Happy to be here. It's it was long due but yeah, finally we made it happen.
Rich Robinson:
Mm, indeed. And always a pleasure. And I would love for you to tell your origin story, your path from your home country to your adopted country of Japan. You are a mathematician, a festival founder, an infrastructure builder, a renaissance man, and I love what you're working on and we'd love to hear your version of your hero's journey, please.
Luis:
Sure.
Well I'm not sure if I could be called Renaissance Man, but okay.
Rich Robinson:
Yes, I just did. I just called you that.
Luis:
But cool. Okay. So yes, I'm originally from Mexico. I was born and raised there. But I've been living all over the world. So I've been living in Germany, Canada, the US, Korea, in Japan but I settled in Japan a while ago. More than sixteen years I think this why we started Mawari in Japan. Indeed, I studied mathematics, the kind of math that is not useful for anything, just to make theory. But I guess now in the world of AI actually that's the one that matters, right?
Rich Robinson:
Who's laughing now? That's actually a beautiful everything's math. So if you're able to really understand the foundation, it's a great building block for everything, right?
Luis:
Yeah. I again like my mom was always very upset because you know in Mexico you need to earn your keep and it's like why didn't you go to full computer science or etc etc. You're not gonna get a good job, you're just gonna be a university professor, which again in a first world country a university professor is well paid and it's a very respectable position, but in Mexico it's more like just you're a teacher in the end, right? And I was like, well but I love math. And so I went into an area that yeah it's just about creating theory but this is how now you put together XR AI and all these abstract concepts into actually something that is manifesting now into reality.
Rich Robinson:
Beautiful, beautiful. I bet I'm sure she's a very proud mom these days now. It worked.
Luis:
I hope so. I mean probably she doesn't understand much of what I what I do.
Rich Robinson:
My mom still thinks I make computers, even though "he works in computers." You know, I'm not sure, but it's okay, mom. It's all good.
Luis:
So that's kind of like my formation, but also since then I've been involved in music, in the creative technology industry per se.
From one side yes I was more on the theory, but on the tech side it was always creative technology. So yes in Japan we co founded digital arts festival. It's called MUTEK. That's actually where the whole idea of Mawari came into fruition because MUTEK by the way, is originally from Montreal. It's I think now this year, this is gonna be the twenty-sixth year. So I actually got to know MUTEK since I was a kid in Mexico, was participant of MUTEK in Mexico. Then yeah in 2015 Takeo and I, our co founder were like something's missing in Japan and it's like yeah, probably something like Mutek would be a very good fit, especially with the way Japan digital art is evolving. So the whole point of MUTEK is to promote and disseminate underground audiovisual art per se. We could call it music as well but in the end it's audiovisual art and most of the artists that perform in MUTEK are like very creative in the sense that they even create their own instruments, their own techniques. It's all about experimenting with music and visuals and using cutting edge technology.
In 2015 is you may remember 2014 is when Meta or back then Facebook bought Oculus and there was kind of like this first bigger wave about virtual reality. So all the artists of MUTEK were starting to get interested into virtual reality and augmented reality as well. But we started seeing a big gap. Like it was not as easy as just, you know, pick up an instrument or do something that as an artist you don't need to go to university to learn how to do it. And in 2016 2017 making or building a VR application, it was not easy. It was very complicated. So the barrier of entry was very high. So this is where Takeo and I started with okay, we need to make a change so that this type of content can be disseminated and distributed faster a way to accelerate this because this creative people is being right now hindered by just the tools that was one of the initial motivations to create Mawari and the second one is that Takeo and I always believed that media will eventually will be all just in 3D. Of like sci-fi movies in the end so and the convergence with AI. So for example, one of the stories that we say in actually in our manifesto is that Takeo was like I want my children to see my mother before she after she passes away and she get to know her as if she was in real life.
And for that you need artificial intelligence and you need extended reality. So it has always been that, like about the that human connection and getting this technology to be something that is normal for every human being. Because also the other thing that we saw is that a lot of people back in in twenty seventeen were horrorized to wear a headset like I look stupid or why I'm wearing this, etcetera. So it was a bit cliche.
Rich Robinson:
Beautiful, beautiful origin story. I think MUTEK is a perfect fit for Japan. If anybody listening has not been to Japan. I was there last year for a month and it is best brand of a country in terms of living up to the experience of being in Japan. And I think bringing that festival there and then having that experience inform the direction forward for Mawari. You know, people often say what people do in their weekend in Silicon Valley, people will be doing in ten years, you know, in their everyday life. And it seems like you really followed something that was your passion and something that was really important for you. And then that opened up this opportunity. And I know in the beginning, as I recall talking to you, did something with KDDI around this whole very serendipitous timing of the AR opportunity. And can you talk about that and that's sort of pull that thread and bring us to where you guys are today.
Luis:
Sure, so yes, Japan is a very unique country. It's has a rich culture, especially Tokyo is cyberpunk in the end literally. So and this influences it is very interesting that you're mentioning about Silicon Valley. This is something that I was actually discussing at a another podcast.
Because Silicon Valley is great in the sense that you know you have this culture of fail fast and recover and until you get to the goal. But that also exactly Japan exactly the opposite.
In Silicon Valley that also gets abused, you know, like the lack of accountability. Like it's okay to raise millions of dollars and never bring a result. Just move on, make your own your next startup. And at some point in time this has been actually praised in the end. One of the requirements to get funded it was that you actually were a serial entrepreneur, whether you had an exit or not, right? In Japan, just for example, with the driving license, is it's not about how many infractions you have. It's a point system that the moment that you have an infraction, you get points deducted. And then it's very hard to recover from that. So it's a culture there where failure is not accepted in the end.
And which is very problematic for startups as well because in the end the VCs in Japan tend to be very strong with the startups and in the end it doesn't work out. But so this I guess is the advantage of Mawari being multicultural in the end. That we have taken the resilience of Japan because we started the company officially in 2017. And this I also applaud Nils and Auki because we are one of the few companies in the AR world that are still out there. And that and we founded a long time ago. Most of the AR companies that were founded in 2017, 2018 is long dead. So that's why this is something that I was talking to Nils the other day. It's like maybe yeah the geography has made us very different. And the fact that yeah we are not locals but in the end we started the companies in a different geography with different cultural values has helped us be resilient and navigate frontier tech because AR in and physical AI well physical AI is really booming right now but it's one of the toughest verticals per se. It's to survive because it takes a lot of time to get there. The Japan thing gave us that unique identity as a company and unique culture.
Now with KDDI. So okay, so back to 2017. So we started the company. We have always had the same mission which is distribute XR experiences worldwide but when it really hit to us is when KDDI actually came to us and to other companies, big companies, startups, they had a very interesting RFP. Is they were already foreseeing in 2018 what's happening right now, embodied AI in the end. It was not called embodied AI, but they came up with okay, so we have a digital human in on Unreal Engine 4. It's powered by Watson. And it's such a powerful application that it cannot run on a smartphone and it cannot run on smart glasses. So we want to stream it. And that was the challenge. But at the time there was no method to stream 3D content. It's just video like what we're streaming right now. So there were some frameworks to kind of fake the 3D with two video streams, for example, two eyes or et cetera, et cetera. But that we realized that would never scale. So in the end we won the RFP. We made our first prototype and we understood a lot. Like also that this type of new content utilize just like AI utilizes a lot of power. This would utilize a lot of bandwidth. But there's a big bottleneck on the networks. We have bottlenecks on GPU consumption, of course, but on networks it's even worse. Is networks are always variable, you never get the same bandwidth. It depends on how many users are into a concentrated area, etcetera etcetera. So there's never the perfect condition for streaming bandwidth. So for example, we can get away with one, two, maybe five megabits per second, but when originally VR streaming was around hundred megabits per second. So you need a very artificial environment to sustain a stream of this kind, up and down.
Rich Robinson:
It sounds like a job for a superhero theoretical mathematician who's also an immigrant founder because immigrant founders have like incredible sort of resilience, what an excellent foundation for a hero's journey. Of course, it's been like the Odyssey of I which I just saw in 70 millimeter at IMAX on film and it's you know 10 year battle and then 10 year return home. But it's been a 10 year battle, but like wow, you guys are kind of like an overnight success after 10 years. Like you did it. Like you actually did it against against all odds.
Luis:
Yeah, it's this is our ninth year, so so that's the whole point. It did take us probably about six years to get the technology to the level that is working right now. And again it was how do we create three D streaming.
I told my our CTO and co-founder Alex is like we need to find a way to create 3D streaming. There was many times that he told me you're crazy why we should be doing that. That's too deep. And I was like but I guess we can do it. We have now I think yeah seven or eight patents already granted on this. And we found a way. Like we became, you know, these R and D partners of Telcos in which they their need was okay, we need to demonstrate that 5G is useful in the end. So and we would be licensing our early stage technology to prove that 5G has a place to be. Because again, yes, you can watch YouTube maybe a slightly better resolution with 5G or in and now 6G is coming into to the picture. But with 6G the use cases, especially with physical AI, are more evident. But 5G it was very difficult to okay, so what is gonna do to my Instagram? What is gonna do to my YouTube or my Netflix? Is 5G really necessary? So there need always needed to be that story that 5G is good for XR streaming, 5G is good for AI processing and all these other narratives. So that helped us because telcos needed to demonstrate this and we needed a sandbox to start testing building our technology, which allowed us to stay unfunded for a very long period of time until we had the technology to a level to say, okay, now it's time to formalize it, build a platform, build a product and platform on top of it. And that's when we ventured into yeah, fundraising, etc. So that's really like the foundation. So we created 3D streaming technology. And just and the other point that we understood with this embodied AI is that 3D streaming is real time rendering in the end. So it's real time processing. So what for example, when you think about the current state of content delivery networks, even this video stream. So the video gets generated either offline and if it gets generated offline, then the CDNs make a copy of this video in a nearby server to the user, right? So this is how it works. Like Akamai Cloud Cloudflare or Fastly etc etc other CDNs cache the content across their globe.
But they're just copying a static file. For live video streaming it works slightly similar, but in the end it's just it's a printed frame. And it's the same frame everywhere, right? But now with 3D you have two things. Okay, so if you have your glasses on, you could be looking at me in a very different angle to what I'm looking at. For example, if we were watching the same thing, you have your own point of view. So a central server that prints the same image, you won't get the it's not working. So you need real time rendering that is personalized and it's
Yes. On top of that also AI. I mean, right now when you talk to your LLMs, it's your own processing unit, right? It's not giving the same answer to you to then to the other person. In that case it was Watson to be connected to the avatar and then talk to you in real time. That and that needed and to feel that it's in real time that needed 30 milliseconds of round trip latency.
So that's what we figure out how to make it happen, but that's what we understood. Okay, so content this new generation of content that is 3D with AI cannot be just distributed as a copy. What we need to distribute is the compute and the application that runs on top of that of that compute. It needs to be lightweight enough because it needs to work at the edge, some of it needs to be offloaded to yes to a bigger server and then it needs to work on the glasses. So this is a concept also that Qualcomm had and that is called a split compute or split rendering. Because depending on the workload and the level of latency that you need, we split the processing of what happens in the experience.
Rich Robinson:
Fascinating, fascinating. I love it. Wow. And it is so relevant to what is happening right now with AI. It's becoming more and more and more relevant and people are more and more and more hungry for this. Yeah.
Luis:
Exactly, because for example, if you ask an avatar or you ask a robot, you expect an answer within less than a second because it's real time interaction. When you are just managing text and Claude or Gemini tells you I'm cooking or whatever he's saying, it's okay. You go on your phone into another application or you do other things and you come back for the result. But real time AI and 3D is very different. It needs to happen within milliseconds. And this is a new challenge and a completely new architecture to distribute this at scale. We understood this and we understood that partnering with a traditional hyperscaler would wouldn't get us there.
We tried actually. We were one of the first partners of AWS on what they called wavelength in this time. The wavelength program was a super low latency edge computing section of AWS. It didn't really scale in the end and this is was kind of like the second light bulb for Mawari is we need to build the right architecture and the network on how to make that happen. So that's where the entire you know community based part comes in. At the time that was now 2021 we were looking into what was happening in blockchain and we saw companies like Helium and Render Network to really scale supply based on crowdsourcing so it was like that's the key because for example when we were trying to do our second raise that was a in the traditional Silicon Valley model, it they were like, okay, get an LOI with a big player that will need your streaming technology and then invest in and put your data centers with GPUs. So I mean that's a big bet. I mean
Yeah, I mean that's the whole point is actually when I pitched the entire vision to Sean, it was very amazing. Like we spent like six hours. A thirty minute it was like a thirty minute pitch and we just started like instead of becoming a pitch because he at the time he was at Borderless Capital. So yeah he was on the busy side. But he's like, no, let's continue talking and we just started brainstorming. It was like yeah sparks like you say instant chemistry. It was magical from that point of view. And yeah he has helped us on yeah like a Helium is not perfect because it was a very early stage. So of course he has given us some guidance on what went well, what went wrong. But the proof is there. If there's enough incentive and people believe in the mission, you can scale so quickly. And that was the key for us because we knew that XR was not there yet to be at the level of mainstream.
We're starting to see product market fit now, but in 2021 there was no product market fit for any XR application. So that's why it would be suicide for us to yeah, to raise a huge amount of money just to buy GPUs that right now with five years down the road would be obsolete. They we couldn't monetize them in this AI era. So we needed to grow organically around pace and this is why also we decided to do it in the decentralized way.
Rich Robinson:
Hmm. Wow. That's a beautiful hero's journey. I love it. And you know, I've read now that you know you're able to take like for instance Unity or Unreal and then cut that bandwidth by four fifths, like a gigantic and then those guys are already optimizing so much, you're able to really with the edge nodes be able to move the needle in a huge way.
Luis:
Yes, okay. So the whole point of our streaming tech is okay, so you can have a heavy scene on the cloud per se on the on the edge. Process it and then we send that in a way that actually you perceive it real time. And that's also the very important 'cause I say it's less than thirty milliseconds. But it needs to fit the requirements of existing networks. So in the end this is why the bandwidth is significantly reduced because we create our own framework to do rendering and streaming in the end because what existed was not cut for this. Because game engines were designed to render on device, not on through a network and not through a distributed system. This is the biggest challenge because for example, let's say mobile gaming or your I don't know Nintendo DS, they work because the processor is wired to the to the display. So there's ultra low latency and you have that instant feedback when when you're playing.
So but just imagine when you when you add a cable you add latency but then you cut the cable and you add kilometers of distance. So what is being rendered and what the user is experiencing, for it to really feel like it's happening in real time is it needs to be around 30 milliseconds of round trip latency. But that's for the image. But if it's interactive and it's for example spatial computing, because this is where we focus spatial computing as well. That and that's part of the split compute, the spatial computing needs to happen on device because if the scene is not anchored sub ten milliseconds your brain will start thinking, okay, this is not really happening here. So there's a lot of technical nuances there because in the end 3D graphics are an illusion because we are seeing in a two D screen something that looks like 3D. It messes with your brain per se, let's call it that way. And your brain has a way to process and perceive information. So XR is all about that, is how you process that information so that it feels like presence. And when you add the processing kilometers away, you need to dissect each one of the components that make that sense of presence real. And optimize for that. So yeah, this is what we've been doing since 2018 pretty much.
Rich Robinson:
Love it. Amazing. And when you really incorporate the whole XR, I know that things are changing so rapidly, but within that experience with motion sickness, between like this presence that you talk about in motion sickness, like how does that all dance together?
Luis:
Okay. So that that's the whole point. More than thirty milliseconds, it gives you motion sickness right away. With the split compute, if some of the things, for example, it could be as simple as like there's some part of the environment that is rendered locally as sub ten milliseconds, then the entire scene in your brain is locking at ten milliseconds. So those are kind of like the tricks that you need to mix to avoid motion sickness.
Rich Robinson:
Got it, got it. Understood. So something that would be annoying just on a two D screen, you know, this is kind of laggy and man, I missed, I got killed and it turns into actual physical sickness with XR. Yeah. Fascinating.
Luis:
Exactly. And any delay makes you throw up, yeah, pretty much. So this is why still all the glasses and all the headsets move from wired to on-site on device processing because it needs to be sub ten milliseconds latency. Otherwise yeah your brain cannot handle it.
It's similar to for example in also in sound, the Doppler effect, you know, the moment that the same sound is split to a certain distance in time, you hear two. But when it's getting closer, you hear it in unison. So it's exactly the same it's a very similar principle. The closer you get to zero, you will never get to zero latency. If we go into to math, is is just a derivative. You will never get to zero.
Rich Robinson:
Never it's always asymptotically. There's gotta be some sort of lag. But the fact is, as you said, you know, I mean thirty milliseconds, it's really difficult for me to even think about that because of course a hundred milliseconds is one tenth of a second. So it's so small, but it's fascinating. Like, you know, I was in Paris for this Raise Summit and I went to you know, the new Notre Dame. And I asked AI about Notre Dame and they look at the church. The church is actually completely asymmetrical. The left tower is a different width than the right tower and these porticos are all different shapes. And they did that deliberately because subconsciously your brain loves asymmetry. And now I can't unsee it, but it's just fascinating about things that really make something very enjoyable and very addictive and very just delightful, are things that we don't even necessarily consciously perceive. Yes.
Luis:
Which is that's one of the points of the digital humans and that's another big topic, the uncanny valley.
Rich Robinson:
The uncanny valley, yes indeed. For our listeners who aren't familiar, the uncanny valley is probably the robot probably looks like me. But unfortunately for the robot, but if it feels too human, but not human enough. But it's this weird kind of like, it feels it creeps me out, right? And yeah.
Luis:
And it's also asymptotically, because you can never get to the human. The whole point is that you can break the Uncanny Valley, but you will never be human. That's the whole point. But we have never we haven't gotten there to the maybe AI video is getting very close to lift that threshold. You can still tell. When a AI video is an AI video and this is not a human person. Although the person generated by the AI model looks agreeable to you, can still feel something itself. But when it's further away is like it feels weird, like it feels very weird.
Rich Robinson:
It's fascinating because I somebody gave me a book. Hey, I wrote a book and then I read the first paragraph and I was like, You didn't write a book. AI wrote a book. I can just tell by the way the sentences are structured and then you see the Will Smith with six fingers eating spaghetti like a weirdo, but then you see something on Grok Imagine now and you're like, Wow, I that's pretty smooth. I guess I guess there is this like sub thirty milliseconds, ten milliseconds. It gets asymptotically to zero, but it gets enough that maybe we do overcome all of these limitations and things become like Westworld that it becomes so hyper realistic that it's indeterminable. I wonder. Yeah.
Luis:
It we will get there. But this is one of the reasons also why we work with avatars right now versus hyper realistic humans because of the uncanny valley.
Rich Robinson:
Yeah, tell us more about your work with avatars. Man, I'm so fascinated. I love what you guys are have built. Wow.
Luis:
That's a Hollywood trick in the end. Okay, so all the animation movies, none of them are they can look hyper realistic. Think about the Avatar movie or any other Pixar movie. All of them are humanoids. But they are not human. That's why your brain accepts them right now. So getting and this is what we learned with KDDI that because they were trying to do a hyper realistic digital human, but we were always hitting the uncanny valley. So that's what we understood. At some point in time, maybe in five years, ten years, we will get to that gap of the uncanny valley that maybe the human will not notice. But we're not there today. And it's many things. It's the facial expressions, the movement of your eyelids, anything that makes you human. So that AI hasn't figured it out yet. I think it will at some point, but
Fascinating. And tell us about your future at Mawari and like what the roadmap is looking like and what you're excited about.
Luis:
We have a platform called Arawa, and we're it's in alpha stage, but it's a content creators platform. So there's this new trend of becoming a virtual YouTuber, which means instead of just being a normal content creator where you show your face, you do it through an avatar. So you don't show your real face. Of course there are a lot of AI girlfriends and other use cases, but this one is more like is you the one puppeting the avatar? Because this is the other thing that we understood. When we were doing the same experience, A B test, a digital human powered by AI about two to three years ago, versus the same avatar but being puppeted by a human. The engagement and the time spent in the experience is was night and day. So for example when you're talking to an LLM you don't feel that personal connection. But when you're talk even when you're talking to an avatar but there's a personal connection, you feel like there's a soul behind that digital presence. You start connecting and you start asking random questions. You start and you forget about the technology. So this is the magic that we have found right now in the product market fit for XR, which is a is digital presence or digital sense of presence. It will get to embodied AI, digital embodied AI at some point in time, but it needs a lot of optimization, just like video models.
Right now we're at the stage of the Will Smith with six fingers and the spaghetti like really not we're not there yet for these autonomous avatars to be chatting with you like you feel nothing is off and you just think of them as your friends. We think the right way to go there is not with digital humans, but for example you can start with the same concept as Tamagotchi. So you can have a digital pet fully AI powered that is more agreeable and little by little then yeah it will evolve into yeah your the Blade Runner fantasy that everybody has. You know the... you know which one I'm talking about.
Rich Robinson:
Yes. Yes, it is. Of course you are based in Japan, so of course, that has to be considered. That's amazing. You're really I think about, you know, the Intercognitive Foundation, our mission is to make the physical world accessible to AI. And I feel like there's these two underground, boring machines that are creating a tunnel and then you're coming from the other side of like you're really making AI accessible to humans. But you're doing it in such a like wow I mean Japan really gives you this incredible space and lab because there is so much you know it's such a large country and it's so technologically advanced and it's so cultural and dynamic the entertainment and the money is there from the big the big companies. So you've been able to do these amazing experiments over and over for the last decade almost and to be able to really uncover things that I don't think there's any other company that's anywhere near close to what you're doing. And it's actually really, really fascinating how you've been able to just, you know, solve problem after problem. It's a beautiful, beautiful journey. And it's just it's like you've been, you know, training for for ten years and now you're really, you know, ten year boot camp. Really a lot of pain and a lot of suffering, but there's all these beautiful things, and now you're like ready to go and do this battle in the AI arena that's so incredible. Tell us and tell our listeners about like where you're looking to like, you know, grow and do new commercial opportunities around the around the multiverse.
Luis:
Well, definitely yes in terms of roadmap, like I say we have the platform Arawa. I think by next year we will be ready to launch it. But we're also very interested in yeah, in collaborating, working on how to make yeah, AI understand the real world but the other way around, like how to bridge that. Because that's also the gap I see with world models because it's only one sided. And they need the other side and I think we can fill that gap.
Rich Robinson:
Absolutely. You guys are the one. Wow. Man, I really would love to come and visit you and see your crazy secret underground labs in Shibuya and have me some Japanese food and soak up some of the atmosphere and spend some time with you. Thank you so much for your time. I love hearing about your journey and really looking forward to reading more about you guys in your glorious future.
Luis:
I appreciate it. And yeah, I mean we are excited for this new phase in general of Frontier Tech and very keen to start collaborating more closely finally with Intercognitive.
Rich Robinson:
Wonderful, wonderful, love it. Thanks once again. Put your hands together for Luis, everybody. That was awesome. Thank you. All the best.
Luis:
Thank you. Cheers.
Rich Robinson:
Adios and sayonara.