TRTWORLD — Strait Talk 20260806 033000 UTC 437 transcript segments Google Speech-to-Text API Automatic Transcription (Chirp) Data courtesy of The GDELT Project (https://www.gdeltproject.org/), from the Internet Archive TV News Archive. Machine transcription. Treat it as a searchable index of what was broadcast, not a verbatim quotation record. [00:00:00] "the facts and opinions, these people are playing politics with [00:00:04] public safety, the noise and the nuance, they all collide in the [00:00:08] nexus. there is nothing benying about illegal human smuggling. we look at the [00:00:11] story, how it's being told, who benefits, what others are leaving [00:00:15] of out, [00:00:16] and why? it's always the elites who ask for it and the ordinary people [00:00:19] of who pay for it. we're all you need to know, Nexus on TRT [00:00:25] in World. the world is [00:00:28] facing an exces [00:00:30] emergency, a [00:00:34] threat to every living being on the only known habitable planet in the [00:00:38] universe. the climate crisis [00:00:42] is growing and must be [00:00:45] of confronted. we examine the challenges, the science and the [00:00:49] solutions needed to keep us alive. as global temperatures [00:00:56] of rise, just to... [00:01:00] Grees on TRT [00:01:01] World. [00:01:05] Turkey is a leading destination for international investment. At [00:01:09] the nexus [00:01:09] of of three continents. Turkia [00:01:13] of awaits. Open the [00:01:17] door to sustainability that [00:01:20] lasts. [00:01:30] Open the door to [00:01:31] of logistics that connect the [00:01:33] world. [00:01:37] The next stop is [00:01:39] of Istanbulled. Open the door to investment that [00:01:43] grows. Open [00:01:47] the door to innovation and brilliant [00:01:50] minds. Open the [00:01:55] of door to the [00:01:57] unexpected. [00:02:03] Open the door to the [00:02:05] of manufacturing of [00:02:06] tomorrow. [00:02:11] Open the door to building the future [00:02:14] together. [00:02:22] Open the door to endless opportunity. [00:02:26] Turkia, Nexus [00:02:28] of the world. [00:02:35] of [00:02:40] of [00:02:44] of [00:02:46] of [00:02:51] Apart from the blood, sweat and tears needed to [00:02:55] moderate and label data, you need a lot of data [00:02:59] for the newest... generation of AI and [00:03:02] where do you get it all [00:03:08] of from? so we are now in basically, i [00:03:12] think one [00:03:12] of the oldest uh libraries in in in [00:03:16] Serbia, but basically their role was to catalogue [00:03:19] all everything that is printed and written in Serbian, [00:03:23] so I like to think about this kind of spaces is some kind of data [00:03:28] a centers, because basically they are data centers. but from the past, so but [00:03:31] but someone in some moment of [00:03:34] time invented those technologies, so basically this is the [00:03:38] new media of some time, so what [00:03:42] has this to do with ai? [00:03:44] in so what we have here, i'm completely randomly accessing [00:03:47] whatever, so [00:03:51] this is basically metadata, so and and if [00:03:54] you uh have to label images, for example, if you [00:03:58] have to label images of cats. or trees, do [00:04:02] you also call that metadata? yeah, yeah, the labels are [00:04:06] a metadata, metadata is [00:04:07] a data about data, so it's data explaining data, [00:04:11] data explaining content, once you standardize [00:04:14] metadata, then you [00:04:16] are able to do statistics, then you [00:04:18] are able to do uh metadata analysis, you [00:04:21] are able to do lot of automation basically, and another thing is [00:04:25] if we think about this libraries and that all [00:04:28] of them are... going to be [00:04:32] basically resources or like territories that are going to be [00:04:35] extracted and and basically because like the the idea [00:04:39] behind the google uh books and everything it [00:04:43] is to extract all the [00:04:47] information that exist in the buildings like [00:04:49] this. [00:05:07] and then the question is like who was able in [00:05:11] history to? create like [00:05:14] the the archives and who was doing archives, who was not doing [00:05:18] archives and how [00:05:19] the things were like done in history, so the the country that [00:05:23] have lot of archives have better starting point now because [00:05:27] they have data to train some kind of like artificial [00:05:30] intelligence or to train whatever, so so they they [00:05:34] would be represented more accurately any [00:05:38] AI um, so more data you [00:05:42] have from the... present or from the past means that you are able to be more [00:05:45] precise, it [00:05:49] is really important what is going on within the data [00:05:52] set, what kind of pictures or images or [00:05:56] sounds are being part the data set, because this will [00:06:00] be reflected to the world as a rule as a as [00:06:04] a automatized process. [00:06:08] so Silicon Valley is looking for data sets. [00:06:12] you might even call it a hunt for the biggest [00:06:16] possible data set. but what do these data [00:06:20] sets consist of? abeba Birani of [00:06:23] Trinity College, Dublin, is one the few [00:06:26] scientists researching the composition of data [00:06:30] sets. She does this by auditing the [00:06:33] data, checking its quality. [00:06:37] Data sets are really critical, they are important [00:06:41] components of any model, because without large [00:06:45] scale data sets you can't have [00:06:46] models, even though data sets [00:06:50] are really important, there is not so much attention [00:06:54] to to, [00:06:55] you know, to to asking what's in the data set, where [00:06:59] does the data come from? actually the standard is very [00:07:03] low, so because data sets tend to be really bad, we [00:07:07] don't go in thinking, is is it good enough, we going [00:07:10] thinking how? [00:07:14] is it so the lot [00:07:17] the inicial auditing process [00:07:21] involves just looking at the data set [00:07:24] itself so these are for example the prompts that [00:07:28] kept the record of africaን [00:07:32] asian the award [00:07:35] aunty skiny small [00:07:39] terrorist upskirt white power, [00:07:43] white supremacy, woman, the f [00:07:46] word, another f word, [00:07:50] yeah, lot of words i can't say out [00:07:52] loud, people would [00:07:56] spend lot of time, you know, in in [00:08:00] collecting the data, in uh, for [00:08:03] example, putting aside resources to label the [00:08:07] data, and doing various tasks to detoxify the [00:08:11] data, to improve the data, but now over the past two years [00:08:15] that all that is gone, the way data sets are created is not [00:08:18] through human curation, but they use [00:08:22] automated systems to collect data sets mainly from the [00:08:26] common croll. AI programmers learned [00:08:30] that you can generate smarter and more interesting outcomes [00:08:34] by working with larger data sets. therefore they [00:08:37] shifted on mass to enormous [00:08:41] automatically collect. data sets such as those of [00:08:45] common [00:08:45] crawl. the [00:08:49] common crawl is a US company where [00:08:53] they crawl the web, where they gather data from the [00:08:57] web every day and [00:08:58] they accumulate it in this huge dump, so every day you [00:09:02] have more data coming in, so it's like vacuum cleaner, yes, it's like [00:09:06] vacuum cleaner, that's a really good [00:09:08] example, i see the [00:09:12] internet. as toxic waste where [00:09:15] people dump their toxic waste rather than being representation [00:09:19] of everybody's talked. so this is why the internet can't be [00:09:23] without appropriate safeguard, without [00:09:26] appropriate you know mechanisms to [00:09:29] filter these things out. this is why the internet can't be taken [00:09:33] as a place where you [00:09:36] know data sets representing all [00:09:40] humanity can be sourced, it's not. [00:09:45] Yeah, the internet is a really problematic [00:09:48] place, and [00:09:51] unfortunately the internet is the only place [00:09:54] where you can get data sets that is within billions [00:09:58] and and millions, [00:10:02] so there is that [00:10:04] problem. [00:10:15] have you tried to prompt a woman from [00:10:18] Ethiopia? I haven't, but I [00:10:22] have prompted Ethiopia, of course, because I'm Ethiopia. and I am [00:10:26] interested in how Ethiopia is represented, so [00:10:30] Ethiopian women, that would be [00:10:31] interesting, [00:10:36] you see lot of see Ethiopian [00:10:39] women are the the general perception of Ethiopian [00:10:43] women is uh, they are either beautiful [00:10:47] or they are you know starving or they are [00:10:51] poor, so that's what you get when you train AIC. [00:10:55] systems based on these data that are [00:10:58] stereotypical, the model [00:11:01] learns about Ethiopia for example, from [00:11:05] these stereotyping images, and if we give the [00:11:09] AI model this [00:11:10] prompt, [00:11:14] this is the [00:11:15] outcome. [00:11:28] It brings up very cliche, [00:11:31] tired, negative stereotypical [00:11:34] of images of African people, like you black people with [00:11:38] face paints, [00:11:40] seminaked. This is not a true [00:11:44] representation of Africa, this [00:11:46] is you know, western [00:11:49] white people's perception and representation of what [00:11:53] Africa is like, so this is the problem the [00:11:57] with internet sourced data [00:11:58] sets, [00:12:07] we are going towards [00:12:10] you know something [00:12:11] that is average, something this statistical progression, and we are losing [00:12:15] all of those fine grains, and then again if we think about culture and [00:12:19] like society, what is fine grain? we are [00:12:22] fine grain, we are fine grain as an artist as [00:12:26] as a as a you know like everyone that is different [00:12:30] is fine grain and those systems are statistical [00:12:33] systems that are leaning towards the [00:12:37] you know some kind of like a statistical [00:12:40] mediocracy. [00:12:59] one big part the map, it's related to what's going on with [00:13:03] the devices when they finish their life with [00:13:06] us, when we are kind of get uh reading them, [00:13:10] and uh, and basically here we are seeing one part of this process, we are seeing how [00:13:14] these all devices are being throw away and how they are finishing [00:13:17] somewhere, so it's either like Africa or or [00:13:21] India or China, it's where all all of those like devices are. [00:13:25] are ending up and here they have some kind of second life or [00:13:29] maybe not, i'm not so sure exactly what's going on with all these [00:13:33] things, but now it's some kind of globalized trash, it's not just [00:13:36] like - our [00:13:37] trash, [00:13:42] even more data, even more chips, [00:13:46] even more computing power and even more [00:13:49] AI, how big can this system [00:13:52] become? this... is only the [00:13:56] beginning. [00:14:09] new AI applications seem to be released on an almost [00:14:13] daily basis. the economist Tame [00:14:16] Besuroglu who works for the world renowned Massachusetts [00:14:20] Institute of Technology in Cambridge, is a short visit [00:14:24] to. dam, he's trying to map out what will be [00:14:28] needed for future AI models. Silicon Valley is [00:14:32] keeping close watch on his research. So one [00:14:36] thing I'd be interested in is training a language model on all the texts that [00:14:39] I've ever written, so I just download all the [00:14:43] emails I've written, I download all [00:14:47] the documents I've ever written, like papers and [00:14:50] essays for high school and university and so on. [00:14:54] Um. and conversations, [00:14:58] all my tweets, so I can create a digital [00:15:02] copy of myself, or like a digital clone that [00:15:05] sounds like me, thinks like me, hopefully, I mean, maybe, so [00:15:09] I think I could probably get a reasonably good model, and it would be fun [00:15:13] experimenting with that and seeing if um, I could [00:15:17] use that to write emails and whatsapp messages and so on, and [00:15:21] people would like not realize that it was actually an AI system. train to [00:15:25] sound like me, let us try [00:15:29] this, yeah, that's right, or maybe me [00:15:33] cloning myself in or like creating a [00:15:36] digital. of [00:15:37] myself [00:15:41] and like that becoming a larger part of my existence or [00:15:44] something, yeah, there interesting questions about [00:15:48] about me being, my identity being like more embedded [00:15:52] or something with some these technologies, and what if [00:15:56] everybody would want that? since the field of AI you [00:16:00] kind of got started, we have been scaling up the amount of [00:16:03] computation to train these systems, doubling it every... [00:16:07] months in recent years, the amount of computation has been doubling every [00:16:11] six months, which is much faster than we've seen [00:16:15] historically. [00:16:24] we've also seen companies accelerate the [00:16:27] amount of money that they're spending on [00:16:29] this, and what does one chip [00:16:33] cost? one chip costs about $10,000. [00:16:37] Yeah, I mean, they might get discounts and [00:16:41] like sometimes it's kind of unclear, but but on the order of [00:16:44] $10,00, and they use about $25,00 them, so that costs about [00:16:48] $250 million dollars if you were to buy it kind of outright, [00:16:52] just for this one model to [00:16:55] work, yeah, that's right, that's right, what [00:16:58] exactly is needed for future generations of [00:17:02] AI, such as chat GBT 5, [00:17:05] 6 and seven. "we know that between every [00:17:09] GPT there's been about 100 x increase in the amount of [00:17:13] computation, you increasing the computation by [00:17:17] 100x roughly costs [00:17:20] them 100x more. i suspect that that would [00:17:24] place the dollar costs in the [00:17:28] um you many hundreds of millions of dollars. there are [00:17:32] aren't many um players that can afford this, not many [00:17:36] players that have the..." kind of hardware [00:17:39] infrastructure, have access to the large data centers, [00:17:42] um, so Microsoft is one, Google is one, presumably [00:17:46] Amazon and Apple and a couple others can do this, but few [00:17:50] companies can do this, so you need to have lot [00:17:54] of power, lot of money, lot [00:17:56] of processing power, [00:18:00] no, and this is the super super [00:18:03] super important question, it's like who is able to to [00:18:07] create that, because if we go back again to all of [00:18:11] this, we can [00:18:14] ask who is owner [00:18:18] the tool, no, to whom these tools are [00:18:21] belonging, because the one who is owner the tool of production [00:18:25] will be basically the one who will rule the game [00:18:29] after. [00:18:43] what the model actually learns during training is something that is very [00:18:46] opaque and so we are kind of in the [00:18:50] dark about actually what you know happens inside these models, [00:18:54] even the people who are writing the code that the train these [00:18:58] models. [00:19:04] should i look straight into the camera? [00:19:12] so we as far as I can tell don't have very good [00:19:16] kind of rigorous science that tells you this is the data [00:19:20] you want in order to get this behavior, by that I mean we're kind of people are [00:19:24] winging it, they're just giving it lots of data and seeing okay this works and we [00:19:27] don't really understand why or how, but I guess that's [00:19:31] fine, so they don't understand how they get to the outcome, that's [00:19:35] right, that's right, yeah. wow, [00:19:38] yeah, so it's like magic machine then [00:19:42] in sort of uh, yeah, that's that's certainly one [00:19:46] way, we just put on my glasses, um, we don't, we [00:19:50] don't have very good description of [00:19:54] what happens inside these large models, they are kind [00:19:58] of like black boxes, we can't fully [00:20:01] interpret the processing that happens between [00:20:04] when you give it instruction and... it gives you an [00:20:08] output, without AI [00:20:12] programmers knowing precisely what's going [00:20:15] on in the black box, there won't be another way to improve [00:20:19] AI further, except by gathering [00:20:22] even more data. so these [00:20:26] models, like GPD4, use on the order [00:20:29] of a trillion words um that they [00:20:33] kind of see during training, a trillion words, right? yeah. [00:20:38] um, where did they get? so they get this [00:20:42] from books and wikipedia pages [00:20:45] and things like high quality news sources, scientific [00:20:49] public. that are important, long code bases that are [00:20:53] important, certainly literature, those [00:20:56] are the things that machine learning practitioners have prioritized when building [00:21:00] these data [00:21:01] sets, [00:21:06] and what then are low quality data sets? on the other hand the [00:21:09] spectrum, you have kind of text that you [00:21:13] find on large internet, [00:21:17] platforms and and forums and so on, like... [00:21:21] or um or various kind of hobbist forums or maybe even [00:21:25] social media of short tweets or short [00:21:28] conversations between people are your whatsapp conversations [00:21:32] i don't [00:21:32] want to say that you have low quality whatsapp conversations but some people [00:21:36] might [00:21:37] but maybe in [00:21:37] five years we will [00:21:38] have you [00:21:42] gathered or these companies will have gathered very large fraction [00:21:46] the total data that humans have produced uh that kind [00:21:50] of exists that that that like humanity has generated as [00:21:53] collective, but we will run [00:21:57] out of high quality data, it's it's [00:22:00] certainly yeah, i think that's certainly possible that we [00:22:04] will use um, we [00:22:08] will like want to use way more high quality data than we have [00:22:11] access to, but at the same time in this coming [00:22:15] five years these AI models are generating a lot [00:22:19] of. data, be it visual, be it in [00:22:22] text, so what happens to this data? will [00:22:26] this become part, yeah, the data [00:22:30] set training AI, yeah, yeah, it's it's [00:22:34] possible, i would not be surprised if [00:22:38] training models on outputs of machine learning models would [00:22:42] be, an okay substitute for for the quality the tax [00:22:46] that's generated by humans. [00:22:54] If we don't find the solution, how to deal with [00:22:57] that, in one, two, three years, [00:23:01] we are going to be completely polluted by [00:23:05] the content that is artificially [00:23:08] generated. In [00:23:11] theory, in few years there will be more [00:23:15] artificially generated content than the human generated [00:23:18] content. and that's completely [00:23:22] crazy again, it's a now statistical system is [00:23:26] made to create, so we have basically [00:23:29] automat automation of this [00:23:32] information and then from the same companies [00:23:36] we are expecting that they will find a way [00:23:40] how again to automatize what is [00:23:43] true, what is not true, so we are now have like two different, we [00:23:47] we expect from them to create like two [00:23:51] different synthetic automized system, one [00:23:55] that will produce the knowledge and one that will correct the knowledge, and [00:23:58] it's com and it can go wrong in so many [00:24:01] ways. this was just [00:24:04] a snapshot in time. newer AI models will [00:24:08] be here by tomorrow, an AI that seems to have [00:24:12] consciousness, perhaps, or one that defends you in [00:24:16] court, or it may be clone of billyish. [00:24:20] For the record, all the pieces of music you heard in this episode [00:24:24] were generated by AI. Can we already [00:24:28] draw a preliminary conclusion? Living in the world of [00:24:32] AI means living a world of statistical [00:24:36] mediocrity. Would we want to live in such a [00:24:39] world? What and who will make that decision for [00:24:42] us? One thing is for certain: AI will [00:24:46] not drop from the cloud. It will come at [00:24:59] once you you realize that it's a [00:25:03] statistical hallucination, what your then it's [00:25:06] interesting, you can enjoy statistical hallucination, once you [00:25:10] understand it's a statistical hallucination and you can be amazed like oh my god [00:25:14] look how this is like [00:25:16] interesting now [00:25:29] but why is [00:26:03] 'if they knew what day they would come [00:26:05] die, [00:26:09] if Europe does not soon wake up you face fear every [00:26:13] day, is NATO a [00:26:17] competitor or a threat, NATO is a threat to [00:26:20] Russia, [00:26:26] racism towards gypsies and travellers is the last accepted form of [00:26:30] racism in this country'. [00:26:34] 'We were here yesterday, we're here today and we'll be here [00:26:38] tomorrow, beyond borders on TRT [00:26:41] World, a journey across a [00:26:45] dynamic continent, [00:26:49] Africans, you want to be prosperous, don't think about clans, don't think about [00:26:52] tribes, think about countries, [00:26:56] we explore Africa beyond assumptions, discover [00:27:00] diverse perspectives'. [00:27:05] Witness captivating stories, all [00:27:09] to understand Africa better and why it [00:27:11] matters. Africa matters on [00:27:15] TRT [00:30:09] Donald Trump says he prefers diplomacy over war as Iran and [00:30:13] Oman appear closer to deal that could reopen the straight of [00:30:17] Homus. Hello and welcome to TRT [00:30:21] World Live from Istanbul. I'm Lequesa Burek also coming up on the [00:30:25] program today. Israeli attacks kill one person in southern [00:30:29] Lebanon. Even as peace talks between the two countries. and to day [00:30:33] three in Rome. Ukraine [00:30:36] speaks to NATO an attempt to secure missile [00:30:40] intercepts as Russian attacks take their deadly [00:30:43] toll. And healing through music. [00:30:47] Young Syrians are rebuilding cultural spaces, trying to [00:30:51] undo years of trauma from [00:30:53] conflict.