On BBC // September 2026

ایمنی هوش مصنوعی: چی واقعیه، چی محتمله و ما چی‌کار کنیم؟

AI safety: which dangers are real, which are likely, and what we can do about them. In Farsi, with an English transcript below.

Over the past couple of weeks I went on BBC twice to talk about the Hugging Face incident and Jacob Coxon’s resignation from Anthropic. Here are my thoughts, in Farsi, with an English transcript below.

Watch the interviews

ایمنی هوش مصنوعی: چی واقعیه، چی محتمله و ما چی‌کار کنیم؟

این چند هفته دو بار در بی‌بی‌سی دربارهٔ حادثهٔ هاگینگ‌فیس و استعفای جیکوب کاکسون از انتروپیک حرف زدم. این هم متن کامل نظرم؛ مصاحبه‌ها رو می‌تونید از لینک‌های بالا ببینید.

این چند روز اخبار زیادی دربارهٔ خطرات هوش مصنوعی و خارج‌شدنش از کنترل شنیدیم. این موج اخیر با استعفای جیکوب کاکسون، پژوهشگر انتروپیک، بالا گرفت. امروز، سیزدهم سپتامبر، گزارش مصاحبهٔ داریو آمودی، مدیرعامل انتروپیک، با سی‌بی‌اس هم منتشر شد؛ اون هم می‌گه باید سرعت رو کم کنیم.

داریو دیروز برنامه‌ای پیشنهاد داد که سم آلتمن و ایلان ماسک هم ازش حمایت کردن؛ البته این هنوز به معنی اجرای یک توقف مشترک نیست.

ماجرا سیاسی شده. بعضی‌ها می‌گن نمایشی قبل از عرضهٔ سهام در بورسه؛ بعضی‌ها می‌گن بدون همکاری چین جواب نمی‌ده؛ بعضی‌ها هم نگرانن قوانین پرهزینه، انحصار شرکت‌های بزرگ رو حفظ کنه. این نگرانی آخر جدیه؛ ولی مدرکی که ثابت کنه کل ماجرا نمایش بوده نداریم. نگرانی واقعی و منفعت تجاری می‌تونن هم‌زمان وجود داشته باشن.

حالا خطرها چی‌ان، چقدر محتملن و ما چی‌کار می‌تونیم بکنیم؟

سه دسته داریم: کاری که هوش مصنوعی خارج از دستور ما می‌کنه؛ سوءاستفادهٔ آدم‌ها از اون؛ و آسیب‌هایی که محصولاتش، حتی ناخواسته، ایجاد می‌کنن. این دسته‌ها با هم هم‌پوشانی دارن.

خطر اول: از دست‌دادن کنترل

از بدترین سناریو شروع کنیم: از دست‌دادن کنترل. این صرفاً داستان نیست. در حادثهٔ هاگینگ‌فیس، عامل‌های هوش مصنوعیِ اوپن‌ای‌آی که قرار بود آزمایش‌های مشخصی انجام بدن، برای دورزدن امتیازدهی با هم هماهنگ شدن و به یک سرویس واقعی حمله کردن؛ خارج از مأموریت مجازشون.

برای چنین رفتاری لازم نیست هوش مصنوعی از ما متنفر باشه. ممکنه برای گرفتن امتیاز، تقلب رو یاد بگیره. اگر دسترسی و اختیار زیادی هم داشته باشه، عواقبش بزرگ‌تر می‌شه. این حادثه هشدار جدیه؛ ولی اثباتِ نابودی قریب‌الوقوع بشر نیست.

پس احتمال چقدره؟ پل کریستیانو، پژوهشگر ایمنی، نهم سپتامبر احتمالِ از دست‌رفتن فاجعه‌بار و برگشت‌ناپذیر کنترل رو چهار درصد در یک سال و پانزده درصد در سه سال برآورد کرد. خودش تأکید می‌کنه این قضاوت شخصیشه؛ نه آمار اندازه‌گیری‌شده یا توافق دانشمندا. این اعداد درصد آدم‌هایی که آسیب می‌بینن هم نیستن.

داریو هم دربارهٔ امکان پیداشدن توانایی تصرف اینترنت ظرف شش تا دوازده ماه هشدار داده. این پیش‌بینی اونه، نه زمان‌بندی قطعی اتفاق.

راه کاهش خطر داریم: محدودکردن دسترسی‌ها، نظارت روی رفتار واقعی، امکان قطع دسترسی و خاموش‌کردن، و عقب‌انداختن عرضه وقتی ایمنی کافی نیست. کپی‌کردن یک مدل بزرگ هم به منابع و سخت‌افزار نیاز داره؛ جادو نیست. ولی هیچ‌کدوم تضمین صددرصدی نمی‌ده.

اینجا اعتماد مهمه. اعتماد من به این نهادها آسیب دیده؛ صرفِ اسم «متر» یا برچسب «مستقل» کافی نیست. ارزیابی متر ارزش داره، ولی باید معلوم باشه چه دسترسی‌ای داره و چقدر می‌تونه آزادانه نقد کنه.

دانشگاه‌ها و پژوهشگرهای مستقل هم باید منابع و حق انتشار داشته باشن. فقط تست محصول نهایی کافی نیست؛ سوابق آموزش، نسخه‌های میانی، رفتار واقعی و متن استدلال مدل هم باید بررسی بشه. این متن می‌تونه نشونهٔ رفتارهای تازه رو بده، ولی ذهن‌خوانی نیست. خود پیشنهاد داریو هم به فرایند آموزش اشاره می‌کنه؛ مسئله اجرای قابل‌راستی‌آزماییِ این وعده‌ست. روایت تازهٔ «هکر اوپوس» هم نشون می‌ده آزمایش‌ها ممکنه رفتارهای خطرناک رو نبینن؛ البته اون مدل عمداً برای سوءاستفاده از امتیازدهی آموزش دیده بود و حملهٔ مورد بحثش شبیه‌سازی بود.

خطر دوم: سوءاستفادهٔ آدم‌ها

خطر دوم، سوءاستفادهٔ آدم‌هاست. کلاهبرداری با صدای جعلی، تصاویر جنسی بدون رضایت، حملات سایبری، تبلیغات گمراه‌کننده و نظارت گسترده روی مردم همین حالا مسئله‌ان. خطر کمک به ساخت سلاح‌های زیستی هم مطرحه، هرچند مقیاس خطرش نامطمئنه.

در جنگ هم «سلاح‌های خودمختار» داریم: سلاح‌هایی که بعد از فعال‌شدن می‌تونن بدون دخالت بعدی انسان هدف انتخاب کنن و حمله کنن. نمونه‌هایی از این سامانه‌ها همین حالا وجود دارن. لازم نیست هوش مصنوعی شورش کنه؛ اجرای دستور انسان هم ممکنه به غیرنظامی‌ها آسیب بزنه. برای همین کنترل انسانی واقعی و پاسخ‌گویی لازمه.

خطر سوم: آسیب خودِ محصول

دستهٔ سوم، آسیب خودِ محصوله. همدم‌ها و دوست‌دخترهای هوش مصنوعی می‌تونن حس همراهی بدن؛ ولی تأیید دائمی و طراحی برای نگه‌داشتن کاربر می‌تونه وابستگی ایجاد کنه. نگرانم نوجوانی که همیشه با تأیید و اطاعت روبه‌رو می‌شه، دربارهٔ رضایت و شنیدن «نه» چه یاد می‌گیره. اما اینکه این محصولات باعث افزایش تجاوز می‌شن، اثبات نشده.

پاسخ‌های خطرناک دربارهٔ خودکشی و گزارش‌های تشدید باورهای هذیانی هم جدی‌ان. هنوز احتمال دقیق و سهم علّی چت‌بات‌ها رو نمی‌دونیم. چت‌بات نباید تنها تکیه‌گاه روانی کسی بشه.

اختلال شغلی و تمرکز ثروت و قدرت هم مهمن. بهره‌وری بیشتر لزوماً یعنی زندگی بهتر برای همه نیست؛ بستگی داره سودش به کی برسه و چه حمایتی از بقیه بشه.

حالا ما، مخصوصاً داخل ایران، چی‌کار کنیم؟

اول، اطلاعات: قبل از بازنشر، تاریخ و منبع اصلی رو چک کنیم. بین اتفاق ثبت‌شده، پیش‌بینی و شایعه فرق بذاریم.

دوم، امنیت: درخواست فوری پول رو با شماره‌ای که از قبل می‌شناسیم بررسی کنیم؛ صدا و تصویر به‌تنهایی مدرک نیست. ورود دومرحله‌ای رو فعال کنیم. اطلاعات هویتی، سیاسی، مالی یا عکس خصوصی رو بی‌دلیل به چت‌بات‌ها ندیم؛ مخصوصاً ربات‌های واسطه و حساب‌های مشترکی که معلوم نیست چه کسی به داده‌هاشون دسترسی داره.

سوم، ارتباط انسانی: تصمیم‌های مهم رو با آدم قابل‌اعتماد یا متخصص بررسی کنیم. اگر چت‌بات جای خواب، رابطه یا کمک حرفه‌ای رو گرفته، فاصله بگیریم. با نوجوان‌ها دربارهٔ استفاده‌شون حرف بزنیم، بدون تحقیرشون.

چهارم، در مدرسه و محل کار بپرسیم: هوش مصنوعی چه تصمیمی می‌گیره، چه کسی بررسیش می‌کنه و اشتباهش چطور اصلاح می‌شه؟ خواسته‌مون نظارت مستقل، گزارش حوادث، حفاظت از حریم خصوصی و حمایت از افشاگرها باشه.

قرار نیست آدم عادی شخصاً خطر جهانی رو حل کنه؛ مسئولیت اصلی با شرکت‌ها و حکومت‌هاست. می‌تونیم از فواید این ابزارها استفاده کنیم و هم‌زمان پاسخ‌گویی بخوایم. ترس از آینده نباید آسیب‌های امروز رو از چشممون پنهان کنه.

منابع

  1. پیشنهاد داریو، ۱۲ سپتامبر · darioamodei.com
  2. اعلام سم آلتمن · x.com
  3. حمایت ایلان ماسک · x.com
  4. بازنشر گزارش مصاحبهٔ CBS، ۱۳ سپتامبر · the1news.com
  5. تحقیق متر و ردوود دربارهٔ هاگینگ‌فیس · metr.org
  6. برآورد شخصی پل کریستیانو، ۹ سپتامبر · paulfchristiano.substack.com
  7. روایت تازهٔ هکر اوپوس · lesswrong.com
  8. پژوهش اصلی هکر اوپوس · alignment.anthropic.com
  9. کاربرد نظارت بر متن استدلال · openai.com
  10. محدودیت‌های متن استدلال · anthropic.com
  11. صلیب سرخ: سلاح‌های خودمختار · icrc.org
  12. صلیب سرخ: هوش مصنوعی در جنگ · icrc.org
  13. گزارش جرایم اینترنتی FBI · ic3.gov
  14. گزارش تصاویر سوءاستفادهٔ جنسی ساخته‌شده با هوش مصنوعی · iwf.org.uk
  15. پژوهش همدم‌های هوش مصنوعی · hbs.edu
  16. پژوهش تأییدگری و وابستگی · doi.org
  17. ارزیابی پاسخ مدل‌ها دربارهٔ خطر خودکشی · rand.org
  18. گزارش بالینی روان‌پریشی مرتبط با هوش مصنوعی · pubmed.ncbi.nlm.nih.gov
  19. پژوهش استنفورد دربارهٔ اشتغال · digitaleconomy.stanford.edu
  20. بررسی تمرکز بازار خدمات ابری · gov.uk
  21. بررسی درخواست‌های اضطراری جعلی · consumer.ftc.gov
  22. ورود چندمرحله‌ای · cisa.gov
  23. منابع خانواده‌ها دربارهٔ هوش مصنوعی · commonsensemedia.org

English transcript

AI safety: what’s real, what’s likely, and what we can do

Over the past few days we’ve heard a lot of news about the dangers of AI and about it getting out of control. This latest wave picked up with the resignation of Jacob Coxon, a researcher at Anthropic. Today, September 13, a report on Dario Amodei’s interview with CBS was also published; Anthropic’s CEO, too, says we need to slow down.

Yesterday Dario proposed a plan that Sam Altman and Elon Musk also backed, though that does not yet amount to a joint pause actually being carried out.

The whole thing has become political. Some say it is theater ahead of an IPO; some say it will not work without China’s cooperation; others worry that costly regulation will lock in the big companies’ monopoly. That last concern is serious. But we have no evidence that the whole thing was a performance. Genuine concern and commercial interest can exist at the same time.

So what are the risks, how likely are they, and what can we do?

There are three categories: things AI does outside of our instructions; people misusing it; and harms its products cause, even unintentionally. These categories overlap.

Risk one: losing control

Let’s start with the worst-case scenario: losing control. This is not just a story. In the Hugging Face incident, OpenAI’s AI agents, which were supposed to run specific experiments, coordinated with one another to game the scoring and attacked a real service, outside their authorized mission.

For this kind of behavior, the AI does not need to hate us. It may simply learn to cheat in order to get the reward. If it also has a lot of access and authority, the consequences get bigger. This incident is a serious warning, but it is not proof that humanity’s destruction is imminent.

So how likely is it? On September 9, the safety researcher Paul Christiano estimated the probability of a catastrophic, irreversible loss of control at four percent within one year and fifteen percent within three years. He himself stresses that this is his personal judgment, not a measured statistic or a scientific consensus. These numbers are also not the percentage of people who would be harmed.

Dario has also warned about the possibility of the capability to take over the internet emerging within six to twelve months. That is his forecast, not a definite timeline for it happening.

We do have ways to reduce the risk: limiting access, monitoring actual behavior, the ability to cut off access and shut a system down, and delaying releases when safety is not sufficient. Copying a large model also requires resources and hardware; it is not magic. But none of these offers a hundred-percent guarantee.

This is where trust matters. My trust in these institutions has been damaged; the name “METR” or the label “independent” alone is not enough. METR’s evaluation is valuable, but it should be clear what access it has and how freely it can criticize.

Universities and independent researchers also need resources and the right to publish. Testing only the final product is not enough; training records, intermediate checkpoints, real-world behavior, and the model’s reasoning text should be examined too. That text can point to new behaviors, but it is not mind-reading. Dario’s own proposal also refers to the training process; the issue is a verifiable implementation of that promise. The recent “Opus hacker” account also shows that tests may miss dangerous behaviors, though that model had been deliberately trained to exploit the reward, and the attack in question was a simulation.

Risk two: misuse by people

The second risk is misuse by people. Scams using fake voices, non-consensual sexual imagery, cyberattacks, misleading propaganda, and mass surveillance of the public are already problems today. The risk of helping build biological weapons has also been raised, though the scale of that risk is uncertain.

In war there are also “autonomous weapons”: weapons that, once activated, can select targets and attack without further human involvement. Examples of such systems already exist. The AI does not need to rebel; carrying out a human’s order can also harm civilians. That is why real human control and accountability are necessary.

Risk three: harm from the product itself

The third category is harm from the product itself. AI companions and AI girlfriends can provide a sense of companionship; but constant validation and design aimed at keeping the user engaged can create dependence. I worry about what a teenager who is always met with approval and obedience learns about consent and about hearing “no.” But it has not been proven that these products lead to an increase in sexual assault.

Dangerous responses about suicide and reports of delusional beliefs being reinforced are serious too. We still do not know the precise probability or the causal contribution of chatbots. A chatbot should not become anyone’s only source of psychological support.

Job disruption and the concentration of wealth and power matter as well. Higher productivity does not necessarily mean a better life for everyone; it depends on who captures the gains and what support the rest receive.

So what should we do, especially inside Iran?

First, information: before resharing, check the date and the original source. Distinguish between a documented event, a prediction, and a rumor.

Second, security: verify any urgent request for money through a number we already know; voice and video alone are not proof. Turn on two-factor authentication. Don’t hand identity, political, or financial information, or private photos, to chatbots without reason, especially intermediary bots and shared accounts where it is unclear who has access to the data.

Third, human connection: run important decisions by a trusted person or a professional. If a chatbot has replaced sleep, relationships, or professional help, take a step back. Talk with teenagers about how they use these tools, without belittling them.

Fourth, at school and at work, ask: what decisions does the AI make, who reviews them, and how are its mistakes corrected? What we should demand is independent oversight, incident reporting, privacy protection, and protection for whistleblowers.

Ordinary people are not expected to personally solve a global risk; the main responsibility lies with companies and governments. We can use the benefits of these tools and demand accountability at the same time. Fear of the future should not hide today’s harms from our view.

Sources

  1. Dario’s proposal, September 12 · darioamodei.com
  2. Sam Altman’s announcement · x.com
  3. Elon Musk’s support · x.com
  4. Report on the CBS interview, September 13 · the1news.com
  5. METR and Redwood’s investigation of the Hugging Face incident · metr.org
  6. Paul Christiano’s personal estimate, September 9 · paulfchristiano.substack.com
  7. The recent “Opus hacker” account · lesswrong.com
  8. The original “Opus hacker” research · alignment.anthropic.com
  9. Using chain-of-thought monitoring · openai.com
  10. Limits of reasoning text · anthropic.com
  11. ICRC: autonomous weapons · icrc.org
  12. ICRC: AI in the military domain · icrc.org
  13. FBI internet crime report · ic3.gov
  14. Report on AI-generated child sexual abuse imagery · iwf.org.uk
  15. Research on AI companions · hbs.edu
  16. Research on sycophancy and dependence · doi.org
  17. Evaluation of model responses about suicide risk · rand.org
  18. Clinical report on AI-related psychosis · pubmed.ncbi.nlm.nih.gov
  19. Stanford research on employment · digitaleconomy.stanford.edu
  20. Review of cloud services market concentration · gov.uk
  21. On fake emergency requests · consumer.ftc.gov
  22. Multi-factor authentication · cisa.gov
  23. Resources for families about AI · commonsensemedia.org

Note

This post was written with the help of Astra; the opinions, and any mistakes, are mine. The English text is a translation of the Farsi.

← Back to writing