معرفی
Vladimir Iashin is a Postdoctoral Research Fellow at the Visual Geometry Group (VGG), University of Oxford. He earned a PhD with distinction in EECS from Tampere University and holds MSc in Applied Mathematics and Computer Science and BSc in Economics from HSE University.
- PhD (with distinction), Tampere University
- MSc (with distinction), HSE University
- BSc, HSE University
His research focuses on deep learning for multi-modal video understanding, particularly audio-visual synchronization and visually guided sound generation. He develops transformer-based architectures for cross-modal sequence modeling and neural audio codecs for efficient spectrogram generation.
His recent publications analyze techniques like codebook-based audio generation and sparse synchronization, with applications in open-domain sound synthesis and automatic evaluation metrics (e.g., Melception, MKL). Key trends include transformer scalability on large datasets (VGGSound) and visual-conditioned sampling trade-offs between fidelity and speed.
Scientific awards include his PhD with distinction. He contributed to open-source tools like Video Features for GPU-accelerated video processing and co-authored benchmarks for container property estimation in robotics.




