Social engineering long relied on a simple technique: making a person trust not a system, but another person. In the past, scammers had to imitate a writing style, choose the right words, and hope the victim would not check the details. Now they have more dangerous tools: artificial voices and faces that can communicate almost like real people. Deepfakes have changed the nature of deception. They no longer look like strange celebrity videos. Today, they can appear as a work call, a relative’s voice on the phone, or a video with a company executive.
Why Old Security Standards No Longer Work
For a long time, security advice remained fairly simple:
- check the sender’s address,
- do not open suspicious links,
- never share a secret code from an SMS.
These rules are still relevant, but they no longer prevent serious problems. When a person sees a familiar face during a video call or hears a colleague’s voice, a completely different level of trust comes into play.
Generative models have learned to copy tone, pauses, and speech patterns. Sometimes, a short fragment from a public interview or a voice message is enough to create a convincing clone. The more active a person is online, the stronger the digital trace they unknowingly leave behind.
When a Voice Becomes a Password
Many people perceive a voice as a form of personal biometrics, replacing a fingerprint. In real communication, however, it no longer always provides protection. A person hears a familiar intonation and automatically fills in the rest of the image. This is where a new kind of social engineering begins. A scammer does not simply ask for money; they create a scene: urgency, fear, authority, or a familiar face. The victim starts acting faster than they can think.

What a New-Generation Attack Looks Like
A modern scheme rarely consists of a single call. More often, it is a chain in which each step pushes the victim toward the desired decision. First, an email arrives in the name of a partner or colleague. Then a message appears in a messenger app. After that, a confirmation call follows. To strengthen the effect, a video is added. A typical scenario may look like this:
- an employee receives an email about the need to send an urgent payment;
- then someone posing as a manager contacts them and asks them not to discuss the task publicly;
- doubts disappear after a short video call;
- the money is sent to an account controlled by criminals.
The person sees several matching signals and stops perceiving the situation as risky. If the director personally joins the video call, why check anything else?
The Case of the Multi-Million-Dollar Video Call
One of the best-known cases of this type of fraud was recorded in Hong Kong in January 2024. An employee of an international company was persuaded to transfer money after a video conference with the chief financial officer and other employees. It later turned out that the video stream had been created using deepfake technology. The damage was estimated at around 25 million dollars.
This case became a warning sign for businesses. Before it, deepfakes were often seen as more of a problem for politicians, celebrities, and the media. The attack showed that criminals had begun to target not only public figures. They need people who have access to payments, contracts, internal systems, and management decisions.
Companies that support remote work have proven especially vulnerable. Video calls have become normal, colleagues may be located in different countries, and urgent financial issues are often resolved without in-person meetings. In such an environment, a fake presence is easier to pass off as part of a regular workflow.

Why Deepfakes Are Hard to Spot by Eye
In the past, fake videos were exposed by mismatched audio and visuals, uneven blinking, and plastic-looking skin. Today, these signs appear less often. Model quality has improved, and people have become used to low-quality video. What once raised suspicion is now easily blamed on a poor internet connection.
With voice, the situation is even more complicated. During an ordinary call, a person does not analyze frequencies and artifacts. They recognize tone, rhythm, and a few characteristic words. If a scammer has studied the victim’s public posts in advance, they can add personal details. The voice clone then becomes part of a broader story. Technological progress is not standing still either. Deepfake detectors are improving, but generators are developing alongside them. They rely on the limits of human attention, especially in stressful situations.
What Needs to Change in Cybersecurity
Organizations should no longer treat voice and video as absolute proof of identity. If a confidential decision is confirmed only by phone, that already creates a risk. Effective measures are not complicated, but they require discipline:
- ban urgent payments without independent confirmation;
- verify unusual requests through a separate communication channel;
- use internal code phrases for critical operations;
- set clear rules for video calls involving financial instructions;
- limit public access to recordings of internal speeches and conferences.
It is especially important to eliminate the cult of urgency. If a process is built so that any top manager can force an employee to transfer a large sum with a single call, then a scammer only has to play that manager.
Ordinary users should review their privacy settings, avoid posting unnecessary voice messages in public chats, and be more careful with children’s videos and family stories. It is even more important to agree with loved ones on verification rules.
Deepfakes do not eliminate trust between people. They force us to build it according to a different scenario. Protection does not begin with complex software, but with the right to pause. A verification call, a backup communication channel, or a code word breaks the attackers’ scenario, which usually relies on the victim making hasty decisions.