Social engineering is no longer just a mass mailing of simple emails. Over the past five years, content generation technologies have changed the very principle of attack: attackers have learned to imitate people so effectively that distinguishing fakes from reality is becoming increasingly difficult.
Deepfake videos, voice synthesis, fake video calls, and mimicked communication styles are now being used against company employees with access to payments, confidential data, and internal systems. While phishing was previously detectable by errors, it is now generated by the same algorithms used in legitimate industries such as film, voiceover, and business automation.
Voice synthesis
Speech synthesis technologies have advanced so much that 10–20 seconds of audio is now enough to fake a voice. Research by Respeecher showed that neural network models can reproduce intonations, pauses, and emotional tone of a voice with up to 95% accuracy.
In the corporate environment, this has led to new attack scenarios: attackers call employees, imitating the voice of an executive, and demand urgent approval of payment or access to the system. Such attacks have already been recorded:
- In 2019, the head of a British energy company transferred €220,000 after a call in which a neural network synthesized the “CEO’s” voice entirely.
- In 2023, in Hong Kong, scammers conducted a video call involving several deepfake employees—all faces and voices were faked.
This is more dangerous than traditional phishing: it plays on urgency, emotion, and hierarchy—people hear a “familiar” voice and act automatically.

Deepfake Videos
Creating realistic videos no longer requires studio tools. Open-source libraries can replace faces in video recordings in minutes, while more powerful models can do so in a fraction of a second. This has led to the emergence of new types of attacks: fake video messages from management, fake instructions, and staged online meetings.
In 2024, the media reported a real case in which an accountant at a large IT company held a video conference with the “director” and several “colleagues.” All participants were deepfake avatars. The company lost over $25 million. The success of such attacks lies not only in visual authenticity but also in psychological factors: people see “familiar faces,” hear familiar intonations, and automatically switch to work mode.
Why traditional security measures don’t work
Antiviruses and filtering systems can’t distinguish a fake call from a genuine one. Deepfake technologies aren’t implemented via a malicious file—they bypass security by affecting the user, not the system. What then becomes a risk factor?
- Employees are used to trusting video calls more than emails;
- Corporate processes often don’t require a second check for verbal requests;
- The brain perceives the visual channel as the most reliable, making the forgery particularly convincing.
The problem is that the human brain is not evolutionarily prepared to encounter an artificially created copy of a familiar person.

How AI Makes Social Engineering Scalable
The key threat of next-generation social engineering is not only realism — it is scalability. Before 2020, attacks involving voice imitation or visual manipulation required technical specialists, manual editing, and hours of preparation. Now the barrier has dropped dramatically: cloud-based AI tools allow attackers to automate deepfake production, generate synthetic voices on demand, and even run targeted phishing campaigns with personalized messages built from publicly available data. Modern attackers combine several technologies at once:
- Automated OSINT collection. Scraping LinkedIn, GitHub, Facebook, and corporate websites retrieves job titles, speech patterns, internal terminology, and communication structure within a company.
- LLM-driven profiling. AI models generate personalized scripts that match the recipient’s writing and conversational style, making messages feel “native.”
- Real-time deepfake generation. Models like VALL-E, ElevenLabs, and open-source voice cloning frameworks produce voice replicas in seconds, without requiring extensive training datasets.
- Synthetic personas. Attackers create entire digital identities — photos, videos, and biographies — to infiltrate corporate Slack, Discord, and Microsoft Teams environments.
A 2024 report by Trend Micro notes that the number of attacks involving synthetic audio increased by over 300% compared to 2021. Meanwhile, deepfake-based fraud costs organizations more than $100 million globally, according to the U.S. Treasury. The combination of speed, realism, and automation transforms social engineering from isolated incidents into mass, structured operations capable of targeting dozens of employees simultaneously.
This shift has important consequences. Companies can no longer rely on intuition or informal communication channels as part of their security perimeter. Any trusted signal — a familiar face on a call, a casual voice remark, a quick message in a messenger — may now be synthetic. Organizations must treat every communication channel as potentially compromised and redesign processes so that no single deepfake is enough to trigger a financial or operational action.
What Companies Are Doing Now: Approaches That Work in 2025
Effective protection isn’t about recognizing deepfake videos, but about changing processes. What could be implemented in such cases?
- Double confirmation of any financial requests—voice and written form;
- Prohibition of urgent transactions via voice command, even if the caller is a manager;
- Clear videoconferencing guidelines (code words, second participant, landline communication);
- Employee training, including real-world attack examples and practice calls;
- Access restrictions—to prevent even a successful attack from opening critical systems.
Some companies use solutions for analysing “unnatural” speech, but they haven’t yet produced consistent results.
The social engineering of the future works not through technical vulnerabilities, but by forging the signals people trust most—voice and visual imagery. No technology can replace critical thinking and a well-thought-out system of internal checks. Companies that adapt their processes to the new reality significantly reduce the risk of attacks, even as deepfakes become more sophisticated and accessible.