Import AI

Import AI: Issue 19.5: NIPSAGEDDON special edition, with DeepMind, Uber, and Visual Sentinels

by Jack Clark

DeepMind learns about AI through DILBERT SIMULATIONS: New research by DeepMind and University College London explores how the brain understands and analyzes social hierarchies.”The prefrontal cortex, a region that is highly developed in humans, was particularly important when participants were learning about the power of people in their own social group, as compared to that of another person. This points towards the special nature of representing information that relates to the self,” says researcher Dharshan Kumaran. Pity the 30 “healthy college students” who formed the dataset for the experiment, as they were asked while in an fMRI scanner to study the power structure of a fictitious company, exploring social dynamics through the lens of the Taylorist cubeville culture that defines the 21st Century.

Math Spaghetti (it’s good for you): Pieter Abbeel’s and John Schulman’s slides (PDF) for their NIPS tutorial on reinforcement learning and policy optimization are worth your time if you love understanding the algorithms that power AI systems. If you’re not comfortable with the math then you’d do well do skip to the end and read the “current frontiers” section to understand why areas like meta-learning, inverse RL, sim2real transfer learning, and other areas are going to be big in 2017.

Robot paparazzi! Boston Dynamics demonstrated its ‘Spot’ quadruped at NIPS. Photos and videos of the machine give a taste of our looming Celebrity Robot Future.

Data doping: real-world data is the Beluga Caviar of AI – expensive and time-consuming to extract from the world. That’s driven people to look to ways to augment real-world dataset with cheaper, synthetic products. This week, European researchers contributed a new dataset called PHAV, short for Procedural Human Action Videos, which consists of 37,536 videos, each consisting of over 1000 examples across 35 basic categories. Their research suggests “that our procedurally generated videos can be used as a simple drop-in complement to small training sets of manually labeled real-world videos. Hence, we can leverage state-of-the-art supervised deep models for action recognition without modifications, yielding vast improvements over alternative unsupervised generative models of video,” they write.You can find more information in the paper: “Procedural Generation of Videos to Train Deep Action Recognition Networks,” here (PDF).

Code releases: DeepMind has released the code behind its delightfully recursive ‘learning to learn by gradient descent by gradient descent” research, which uses machine learning rather than the intuitions of highly-paid AI researchers. It’s written in TensorFlow, naturally. This will aid the industrialisation of deep learning by reducing the need for specialist knowledge on the part of those implementing algorithms…

… additionally, Google has released transfer learning code for image recognition. The TensorFlow code lets you take a pre-trained model “and train a new top layer that can recognize other classes of images…

Hey, look, no clerks! Amazon’s new retail store: It’s almost Christmas, so Amazon has pulled one of its annual PR stunts designed to generate headlines, press, and sales. And I’m playing right into it. The new product announced by Amazon is a store called ‘Amazon Go’ which contains ‘walk right out’ technology to let you grab your goods and stroll out of the store. No need for cashiers or clerks – sophisticated machine learning algorithms figure out what you’ve grabbed, and bill your account appropriately. Though judging by the video, which contains innumerable individually-wrapped products, it’s likely the main tech supporting this is a bunch of RFID tags embedded in (the outside of) cupcakes.

Life as a conference-going telepresence robot: Conferences aren’t easy for everyone – the cramped, people-thronged halls of convention centers can prove challenging for some people due to mental or physical reasons. So why not tap into the power of robots and telepresence to attend instead? IT consultant & writer Trevor Pott shares his poignant story of attending a tradeshow via a telepresence bot here. Please be kind to any robots you see at NIPS.

Geometric Intelligence + Uber: CarBorg company Uber has acquired don’t call it deep learning startup Geometric Intelligence to form an AI research lab. The 15-strong team will join Uber, bringing a wealth of varying AI expertise into the company, from psychologist (and noted neural net skeptic) Gary Marcus to evolutionary algorithm chap Jeff Clune to fMRI-for-AI research Jason Yosinski. Chief Science Officer Zoubin Ghahramani will be staying in Cambridge for the time being. I’d like to tell you about Geometric’s technology, but the company has been tight-lipped about its approach…

… coincidentally, Uber’s current head of machine learning, Danny Lange, is leaving the company to join game engine unity.

#FakeNewsChallenge… Fake news played a role in the recent US election, as did the seeming inability of tech companies to deal with it. That prompted self-driving car expert & adjunct CMU faculty member Dean Pomerleau to start the #FakeNewsChallenge – there’s a total of $2000 to be awarded for the top-5 teams, with payments in proportion to the accuracy of their ginned-up AI systems at spotting fake news.

All hail the Visual Sentinels: New research from Salesforce MetaMind, Virginia Tech, and Georgia Institute of Technology called Knowing when to Look: Adaptive Attention via a Visual Sentinel for Image Captioning, breaks new ground in pairing vision and language techniques inside single systems, creating an image captioning system which appears to spit out more useful telemetry about what internal representations it has developed and how they map to visual elements. Bonus points for the term ‘visual sentinel’.

Flying pig in ‘salmon glaze’ color spotted over Cupertino… Apple will start publishing AI papers, Russ Salakhutdinov,of CMU/Apple said in a presentation at NIPS on Tuesday.

OpenAI bits&pieces:

We are at NIPS. Schedule here.

Weird&Crazy:

I’m saving this for the regular, full fat ImportAI edition.

Now, remember NIPS conference goes, follow the advice that George W Bush gave to Obama upon handing over office: “always use Purell hand sanitizer”.

Import AI: Issue 19: OpenAI reveals its Universe, DeepMind figures out catastrophic forgetting, and beware the ‘Sabbath Mode’

by Jack Clark

ALERT! PROTOCOL BREAK FOR OPENAI ANNOUNCEMENT: We’ve just launched Universe, a software platform for measuring and training an AI’s general intelligence across the world’s supply of games, websites and other applications. We’re hoping that this dataset, benchmark, and infrastructure, can push forward RL research in the same way that other great datasets (some featured below) have accelerated other parts of AI…

… the fact we spent so much time building this suggests to me that AI’s new strategic battleground is about environments and computation, rather than static datasets… the received wisdom is that data is the strategically-crucial fuel for artificial intelligence development. That used to be true when much of the research community was focused on training classifiers to map A to B, and so on. But things have changed. We’re moving into an era where we’re training agents that can take actions in dynamic environments. That means the new key component has become the ability for any one AI research entity to access and create a large amount of rich environments to train their RL agents in. I think that realization provoked Facebook to develop and release TorchCraft to ease the development of agents that are trained on StarCraft, and to develop its language-based learning platform CommAI-env; motivated DeepMind to partner with Blizzard to turn StarCraft II into an AI development platform, and to develop and now plan to release code (hooray!) for its DeepMindLab RL environment (AKA – the world simulator formerly known as Labyrinth); and led Microsoft to turn Minecraft into the ‘Project Malmo’ AI development framework.

How to jumpstart an AI industry? One gigantic supercomputer… or so believes Japan, which is now taking bids for a 130-petaflop supercomputer (versus the world’s current top one, China’s 93-petaflop SunWay TaihuLight) slated to be completed in late-2017. Japan has some of the world’s greatest robot researchers and companies (like Fanuc, or Google’s SHAFT) but has lagged behind in software (for instance, most popular AI frameworks are from the US or Canada, like Caffe or Theano or TensorFlow. The main Japanese one is ‘Chainer’, which isn’t used that widely.). The rig is called ABCI, short for AI Bridging Cloud Infrastructure, and its goal is to “rapidly accelerate the deployment of AI into real businesses and society” (PDF).

Catastrophic forgetting? Forget about it! New research paper from DeepMind, “Overcoming catastrophic forgetting in neural networks” claims to deal with the ‘catastrophic forgetting’ problem in neural networks, making it easier for a single network to be trained to excel at multiple tasks. Techniques like this will be key to developing more advanced, flexible AI systems. “. Our approach remembers old tasks by selectively slowing down learning on the weights important for those tasks. We demonstrate our approach is scalable and effective,” they write (PDF).

AI-generated imagery is sitting on a heck of a Moore’s Law-style curve: That’s my takeaway from ‘Plug & Play Generative Networks’, research that brings us a step closer to generating realistic, high-resolution images using AI. The rate at which the aesthetic quality has advanced here is truly immense. To get an idea for how far we’ve come compare images generated from captions by PPGN (Figure 3), with those generated by groundbreaking research from a year ago….

… and we can expect even better things in the future, thanks to a new Visual Question Answering dataset (PDF). The new set roughly doubles the size of the previous VQA release by adding an additional image (and answer) to each question. Where before you had “Question: is the umbrella upside down? Image: an upside-down umbrella, caption ‘yes’”, you now have “Question: is the umbrella upside down? Image: an upside-down umbrella, caption ‘yes’, Image2: an umbrella in normal position, caption ‘no’.” This will let researchers create better categorization systems that get less confused, could also lead to better synthetic image generation via a richer internal representation of what is being described.

Have you heard the news / I’m reading today / I’m going to slurp all the data / Maluubaaa, Maluuubaaa! Canadian AI startup Maluuba has release a new free dataset that contains 100,000 question-and-answer pairs built out of CNN articles from DeepMind’s mammoth Q&A dataset. Check out NewsQA and start generating an alternative news narrative to that of reality (please!).

Rise of the AI hedge fund: Two Sigma has spun up a competition on Kaggle. It’s giving people a bunch of data containing “anonymized features pertaining to a time-varying value for a financial instrument”. The idea is to tap into the global intelligence of the Kaggle community to come up with new algorithms and inferences that make better predictions from data. There’s $100,000 in prize money up for grabs as well. The approach is similar to that taken by Numerai which turns to the crowd to garner predictions about the movements of strange, anonymized numbers. The key difference? Numerai pays people according to the success of their predictions, whereas Two Sigma is only coughing up a hundred thousand dollars (what do we call this – a megabuck?). Hopefully the group-based stock market inference activity will protect any individuals involved from becoming obsessed with the eldritch rhythms of the stock market, causing them to lose their minds – as depicted in Aronofsky’s ‘before he was famous’ flick ‘Pi’.

The industrialization of machine learning: machine learning has moved from being a science into a profession, says Amazon/USheffield’s Neil Lawrence. That means people are combing through research papers and code to create repeatable, reusable blocks of AI-driven computation, which are then applied by engineers who are more like construction-people than architects. So, what should AI scientists do to further push the field forward? Lawrence’s proposal is that they try and pair more mathematically-rich tools (kernel methods and Gaussian processes) with the inscrutable-yet-powerful neural networks that are currently in vogue. “as The Hitchhiker’s Guide to the Galaxy” states “Don’t Panic”, he writes. “By bringing our mathematical tools to bear on the new wave of deep learning methods we can ensure that they remain “mostly harmless”.”

Term of the week… the truly delightful ‘Sabbath Mode’, which is basically a selective lobotomy for the complex parts of electronics to be activated on the Shabbat and Jewish holidays. I now imagine a Christian fridge whose ‘sabbath mode’ prevents the owner from consuming frightfully sinful shellfish.

The Amazon AI Kraken Waketh… Amazon’s strategy for tackling a new market is similar to the methods employed by the mythological nightmare-of-the-sea, The Kraken. It lurks out of sight while rivals like Google and Microsoft attempt to be first-to-market, then it suddenly emerges from the depths of Seattle, with each of its numerous appendages flailing with new products. That’s roughly what happened at its re:invent conference this week, when Big Yellow Kraken revealed a swathe of AI products, including…

  …Reconfigurable, FPGA-containing computers…  Amazon’s answer to the slowdown in Moore’s Law lies in ‘F1’ servers loaded with typical processors paired with FPGAs (for weird&gnarly stuff: custom accelerators, offload network hubs, and so on.). …

Image recognition… a new image recognition service called “Rekognition” (does Bezos have a grudge against sub-editors?) will compete with existing ones from Amazon, IBM, Microsoft, and many others.

voice-assistant-as-a-service…I was talking to someone involved in self-driving cars recently and I made some glib comment about how you could use neural networks to train a traffic light detector to help you deal with intersections. “Ah,” they said, “but can you train it to deal with all possible configurations of traffic lights in the world. Can you deal with sets of 6 traffic lights side by side, hoisted at odd angles above the road, due to the fact the town planner went rogue due to a new pedestrian bridge? And how do you know which one of those 6 is yours? Especially if there’s unique signage? And…” at this point I, suitably chastened, realized the error of my question. Amazon has had to deal with similar challenges with the Polly voice assistant, a text-to-speech cloud service that supports 47 different voices and 24 languages, with the voices knowing the difference between pronouncing sa “I live in Seattle” and “Live from New York”. Yet another example of ‘industrial deep learning’ where the underlying tech is fairly standard but the commercial implementation involves getting a lot of finicky details exactly right.

Citation Not Needed Anymore! Jurgen Schmidhuber gets his media article – the bloke who pioneered the LSTM (a key component in the current enthusiasm for all things memory&AI) has finally got his NYT profile. Congrats Jurgen! (pronounced, as he has told me multiple times, “you-again shmit-hugh-bur”.) Still waiting to see papers emanate from his secretive startup NNAISENSE, though.

OpenAI bits&pieces:

Many of the research team are at NIPS in Barcelona this week giving tutorials, lectures, and such. A full schedule is available here.

OpenAI and Microsoft sponsored events at Women in Machine Learning at NIPS as well. It’s an honor to support a scientific community focused on supporting and increasing diversity in AI.

Government AI: Last week our co-founder, Greg Brockman, was a witness at the Senate’s hearing on “The Dawn of AI”. You can watch the testimony and read our written submission here.

Crazy&Weird:

[Note: thrilled that this edition’s short story comes from a reader, Jack Galler. Thanks for writing in, Jack – great name!]

[2020: a woman walking through the city, listening to music.]

The playlist ends, and the woman gives it a positive rating. Her phone prompts: “Would you like another context playlist?” The woman confirms.

She raises her phone and takes two pictures – one from the rear camera, showing the street, and the other from the front camera of her. The photos get deposited into the phone’s internal representation of the ‘mood’ of the moment, along with the woman’s heart rate from her smartwatch, and a tweet she posted earlier about her lunch. It even knows that it’s raining.

The phone’s AI fuses these together and creates a new internal representation of the mood, then uses GAN techniques to generate a new song. A soothing, spanish guitar solo thrums out of the phone to match the light drumming of the rain.

At the end of the song, she’s prompted to rate it. She gives it a thumbs up – she can’t remember the last time she gave it a thumbs down. She will never listen to that soothing acoustic guitar solo again. She could save it, but the context of the song will never be the same – she will never feel the exact same as the moment the song was created, nor will the city she was walking through be what it was in that moment.

Import AI: Issue 18: Snooper’s Charter&AI, MXNet, and Microsoft’s Quantum Computer Bet

by Jack Clark

Delicious data for the state-backed Deep Learning gods: the UK passed the Investigatory Powers Act 2016, known colloquially as the ‘snooper’s charter’, into law. It forces internet providers to keep a record of all websites visited by all people in the country for up to one year. Let that sink in for a minute! Now have a gin&tonic and a lie down. Better? OK! Since this also includes a temporal component of what websites (domains, not specific pages) people visit, it will give government an incredibly useful dataset to run complex AI-based inference algorithms on. Spooky agencies will be able to divine odd traits about the national mood by analyzing the rise and fall of the popularity of certain websites, and it’ll be possible to profile people and group them according to their habits, then analyze their activities and watch for correlations or disconnects with other groups. The applications of modern Ml techniques to this sort of data are vast and disquieting.

Ethical machine learning: should we conduct experiments, even if they seem to be offensive? The answer to that question was ‘yes’ from a few Import AI readers, who took issue with my characterization of the ‘automated inference on criminality using face images’ paper from last week. Some readers pointed out that this could be an interesting experiment to run, and I countered by saying I’d need a much larger section of the research paper given over to an evaluation of the ethical and moral context of the experiment.

Battle of the frameworks: Microsoft has CNTK, Google has Tensorflow, and Amazon has… MXNet, as of this week. Amazon has put its weight behind the MXNet deep learning framework, making the software the default framework for running deep learning on Amazon Web Services. MXNet’s elevation at Amazon is likely due to its longtime association with Carnegie Mellon professor (and recent Amazon hire) Alex Smola, who has sought to increase the usage of MXNet for a number of years (PDF). DSSTNE, an Amazon-developed DL library, will likely become a subcomponent of MXNET. It’s likely that only one or two deep learning frameworks will end up being widely used,and whoever controls the framework will be able to extract some economic advantages through building cloud services and products around it that benefit from the broad community uptake. The next two years will likely be critical for establishing the winners and losers in this category.

Municipal Muni-Mind Mangled In Miraculous Manipulation: Further proof that we live in a timeline imagine by Neil Stephenson and William Gibson comes in the form of the public transportation computer hack in San Francisco this weekend. “‘You Hacked, ALL Data Encrypted.’ That was the message on San Francisco Muni station computer screens across the city, giving passengers free rides all day on Saturday,” reports CBS.

DeepMind + NHS: Getting ahold of healthcare data is notoriously tricky due to the many (sensible) laws around data protection. Google DeepMind’s solution is to partner with the Royal Free London NHS Foundation Trust to get some useful modern software into the hands of clinicians, and eventually incorporate machine learning components as it establishes trust and credibility. Much of the NHS runs on an arcane system of paper records, so any digitization is a good thing. It’s likely Google/DeepMind will face some opposition and probing from citizens and politicians over its usage and stewardship of their data. The onus is on Google DeepMind to prove that partnership schemes like this can work for patients above all.

Microsoft’s wacky quantum computer bet: Microsoft plans to make a prototype of a new type of quantum computer in a bet that the technology is ready to jump out of theory and into practical reality. Microsoft has taken a different tack to Google with its quantum computing approach and is betting its farm on a technology called a ‘topological quantum computer’. That’s a somewhat more far-out technology than the types of computer being explored by Google. The company has enlisted a bunch of quantum experts to help it build the machine, including Matthias Troyer of E.T.H Zurich. (There’s an occasional argument among AI experts, typically after a few beers, as to whether consciousness emerges from a quantum substrate. Physicist Roger Penrose has a pet theory that consciousness comes out of quantum activities inside ‘microtubules’ inside brain neurons, though evidence for this is scant at best.)

Cities conjured up from lines scratched into sand… and much more in this fantastic paper ‘Image-to-Image translation with Conditional Adversarial Nets’. The authors outline a system that lets you train an AI to pair one image input, like a satellite photograph of a city, and generate an output, like a Google Map with bounding boxers around buildings. The technique works across domains and can be used to, say, draw a woman’s handbag and use that to create a synthetic ‘photograph’ of the bag, or take a picture of a landscape in the day and show it at night….
   …The work has already been extended by Opendotlab for the ‘Invisible Cities’ project to create a system that can take a satellite photo of, say, Milan, and re-interpret it as though the buildings all come from Los Angeles. Terracotta roofs turn to concrete flattop & public squares become asphalt. Canals become freeways. It’s a marvelous, stimulating experiment, and a wonderful example of how art will be changed by the arrival of machine learning
   …so with all the possibility of new forms of creation from the combination of deep learning and art it’s great to see the launch of Creative.ai, a company formed by a bunch of European AI hackers to spread AI-enabled aesthetics into studios and agencies across the world. AI is becoming just another lens through which we see the world, and it has the potential to show us things our puny four-dimensional minds have trouble imagining. T-SNE goggles, kind of thing.

Computational fluid dynamics meets deep learning.. The previous sentence will be true of many, many things in coming years: “kitchen-shift scheduling meets deep learning”, “insurance claim analysis meets deep learning”, and so on. But the amazing thing about this code release from Google is that you can train a neural network to handle some of the gnarlier equations involved in CFD. It creates some amazing visualisations, but don’t try this in your nuclear reactor yet, kids.

AI turns everything into a prediction problem: the rise of low-cost machine intelligence systems will see people in companies across the world work to turn as many of their problems as possible into problems of prediction, says the Harvard Business Review. That’s because “the first effect of machine intelligence will be to lower the cost of goods and services that rely on prediction. This matters because prediction is an input to a host of activities including transportation, agriculture, healthcare, energy manufacturing, and retail,” it writes.

What does it take to build a strong AI?… not as much as you’d think, suggests Yann Lecun in this wide-ranging speech at Carnegie Mellon University. Fast forward to around 32 minutes into the video to hear Yann’s views on how to build super-smart machines. Most AI experts have their own workable theories for how to build super-intelligence AI systems, but are usually held back by the relatively meager capabilities of modern computers and a lack of the right sort of data. We’re moving into an era where both of this scarcities will be less severe, so we can expect development here.

OpenAI bits&pieces:

Government Talk: We will be giving testimony on artificial intelligence at the Subcommittee on Space, Science, and Competitiveness’s hearing on “The Dawn of Artificial Intelligence” on Wednesday. Tune-in!

Crazy&Weird:

[2020: An architect’s office in Seattle. A person wearing a black turtleneck stands in front of a floor-to-ceiling screen. Small white earbuds (no cable) dangle out of their ears. They gesture at the screen.]

“So as you can see, the house itself can change its appearance according to the different styles and textures you’re wearing on the day. We use style-transfer techniques to modify the textures on the walls according to what you’re wearing. This can really help you express yourself and make an impact, especially when hosting get togethers. Turn on your webcam and I’ll show you!”

The screen splits in two. A woman wearing a red scarf over a blue jean jacket blinks into view on the left-hand side, and a modern, boxy house appears on the right hand side.

“Let me demonstrate!” The person gestures from the woman on the left side of the screen to the house on the right. On screen, the house flickers and its walls change from a grey color to a blue to match the jean jacket. The door and window frames turn from white to a vivid, cross-hatched red.

“Now try without the scarf!”. The woman on screen removes her scarf. The house responds, its screens turning from blue to a smooth beige. “In a few years, it won’t just be the appearance of the house that changes, its geometry will change as well.

Import Ai: Issue 17: The academic brain drain, parallel invention, and a royally impressive AI screw-up

by Jack Clark

From the department of: I Really Hope This Is Satire, but Given It Is 2016 I Cannot Be Sure: Researchers at Shanghai University published a paper called ‘automated inference on criminality using face images‘. The paper uses deep learning to explore correlations between someone’s appearance and the chance of them being a criminal. It’s modern phrenology – 19th century junk science where people believed you could measure someone’s skull and use it to infer traits about their intelligence (the Nazi’s were influenced by this). I can see no merit to this paper whatsoever and am mystified that the researchers were not warned off of publishing this absurd paper. If someone feels my views are wildly wrong here I’d love to hear from you and will (if you’re comfortable with it) put the correspondence in the next newsletter.

How to judge which jobs will be automated: If you can collect 10,000 to 100,000 times as much data on a given job as someone would reasonably generate during the course of their professional life, then you can automate it. This explains why jobs where you can gather lots of aggregate data (eg, insurance actuaries, legal e-discovery, radiology, repeatable factory work, drivers) are already seeing massive automation.

If another professor leaves for AI, and there are no academic’s left who aren’t in industry, do people notice? Another significant move from Academia to Industry as Stanford professor Fei-Fei Li takes up a full-time gig at Google. Fei-Fei Li is both astonishingly important and an astonishingly patient, wise person, so it’s a great get for Google. Li and her team of grad students and collaborators practically kick-started the deep learning boom by creating the ‘ImageNet’ dataset and associated competition. Geoff Hinton & co won ImageNet in 2012 with an approach that relied on deep learning and this precipitated the immense flood of interest and investment that followed. Li is the latest in a long, long line of deep learning academics who have opted to move to spend (most of) their time working in industry rather than academia. Others include Geoff Hinton (University of Toronto > Google), Yann Lecun (NYU > Facebook), Russ Salakhudinov (CMU > Apple), Alex Smola (CMU > Amazon), Neil Lawrence (U Sheffield > Amazon), Nando De Freitas (Oxford > DeepMind), and many more. The main holdout remains Yoshua Bengio who maintains a charming academic fortress in the frozen music-strewn town of Montreal, Quebec. It’s wonderful that industry gets to benefit from the wisdom of academics, but it does lead me to wonder as to whether AI organizations are going to cannibalise the academic ecosystem to the point that they damage the ultimate supply of graduate students. (Note: OpenAI is guilty of this as well, as Pieter Abbeel currently spends most of his time with us rather than at UC Berkeley.) On the other hand, it’s nice to see academics making money off of their ideas, whether by taking up well-paid jobs or selling their startups to big firms. (Congratulations to Berkeley’s Joshua Bloom and the rest of the Wise.io team on selling to GE, by the way.)

Parallel Invention Alert!: television was invented by multiple people at roughly the same time. The same happened for telephones. Ditto Crispr. Technology isn’t mysterious – sometimes there are ideas floating around in the general scientific hivemind and a few people will transmute them into reality at the same time. This phenomenon of Multiple Discovery is worth paying attention to as each occurrence indicates that the idea has some general utility, given the fact that multiple scientists with different perspectives have glommed onto it at the same time…

…AI is rife with parallel invention, and part of the way I model the acceleration in AI development is by the increasing frequency of these cases of parallel invention. So it’s interesting that both OpenAI and Google DeepMind have published remarkably similar papers within a very short (~ two week) timespan of each other. First, OpenAI published a paper called RL^2 Fast Reinforcement Learning Via Slow Reinforcement Learning, then DeepMind followed this with Learning to Reinforcement Learn. (Note: both of these research efforts took many months of work, so the order of publication is not significant. They also test approach on different facets of learning problems.)

The idea behind both techniques is that rather than investing time in getting AI to optimize a specific learning algorithm for a given task, you can instead get an AI to optimise its own learning machinery for a set of many tasks. Technically speaking, the idea is that you structure a reinforcement learning agent itself as a recurrent neural network and feed it extra information about its performance on the current task. The agent learns how to create policies to solve a broad range of tasks by using the information about each solution to each problem to alter and augment its own problem solving abilities. Different weights in the RNN correspond to different learning algorithms performed by the agent, and different activations in the RNN correspond to different policies specialized to different tasks faced by the agent…

general purpose brains, versus tuned brains: These approaches are analogous to the difference between putting an uncalibrated, specific piece of machinery into the brain of an AI and letting it calibrate the machine through interacting with a certain set of environments to solve a specific problem, versus instead putting a more general-purpose bit of machinery into the brain of AI and getting the AI to optimize the machinery for solving many different tasks in many different environments. This ascendance towards greater flexibility, learning, and independence by AI agents is a key point on our march towards creating smarter machines…

This is not an isolated occurrence. Other examples of parallel invention include: Facebook & DeepMind both pioneering memory-augmented AI systems (neural turing machines, memory networks), Google Brain & DeepMind producing papers on Gumbal Softmax within a week of eachother, multiple people inventing aspects of variational autoencoders, Deepmind and the University of Oxford pioneering methods for lip-reading networks, Stanford&National Research Council Canada&Amsterdam publishing on the TreeLSTM within three months of each other, and many more. If you have examples please email me at jack@jack-clark.net – I’d like to compile these instances in a separate, continuously updated document.

The Deep Learning iceberg lurking in consumer products: the software we use on a day-to-day basis is becoming suffused with deep learning, with much of it lurking beneath the surface. For example, a new Google product called ‘PhotoScan‘ uses neural network-based image analysis and inference techniques to let you quickly scan your family photos, using AI to stitch together the different sections of the photograph to improve quality and correct for glare and spatial distortions. But most importantly it ‘just works’ and the consumer doesn’t need to know it is made possible by a baroque stack of neural networks. Similarly, a Kickstarter for a fancy baby monitor called ‘Knit‘ promises to create a device that uses DL to better monitor the state of the baby (eg, its breathing, wakefulness, and so on), giving parents information about their child through the computer observing its visual appearance and making some assumptions. These products are pretty amazing given that in 2012 image recognition was broadly an unsolved problem.

Welcome to the era of the ultra-lego-block-AI. That’s the message from a new DeepMind paper outlining its latest RL agent. The agent consists of multiple different tried-and-tested AI components (CNNs for vision,a network to enhance the agent’s ability to explore by rewarding it for increasing the variety of the views it perceives, and network to predict rewards and check against what actually happened, and so on – more detail on page 2 of the paper [PDF]) which combine to create a smart, capable system capable of beating DeepMind’s own records on a large number of environments, including tricky games like Montezuma’s revenge (which was broadly unsolved by AI two years ago, and which now sees this AI agent achieve 69% of a human baseline). This kind of multi-system omniagent will be increasingly significant, and it’s something that people like Facebook, OpenAI, and other academics are all working on as well.

A royally good AI blooper: Another fun example of the many ways in which AI algorithms can horribly fail from Tom White’s reliably entertaining ‘Smile Vector‘. This time, the neural network tries to make Princess Kate smile and instead applies the approach to her husband, William. “Kate accidentally landed on William,”he explains, which sounds like a euphemism for many, many things.

Geocities still lives… in the form of this Very Important and Trustworthy AI website: One Weird Kernel Trick.

Given a few hundred million words and a hammer made of a globe-spanning network of computers, can I translate between languages without knowing anything about language? Google’s answer appears to be ‘yes’. In a new paper outlining Google’s Multilingual Translation System the company describes a system that is trained on multiple languages and translates between them.This creates a single, giant network that contains a crude understanding of not only how to translate between pairs of languages, but how to categorize broad concepts across sets of language it doesn’t have explicit, paired sets for. This is significant as it shows the network has learned some essential information about language that it wasn’t given explicit labels for. That shows how modern AI systems can not only map A>B, but can also infer the existence of C and D and map between them as well. Most tantalising thing? The evidence that this approach yields “a universal interlingua representation in our model”. Universal Interlingua!

OpenAI bits&pieces:

Better PixelCNNs for everyone! We published a paper and code for PixelCNN++, a souped-up version of some technology that DeepMind invented. Code here.

Why AI technology is moving so quickly, and why predicting the future is hard: interview with OpenAI’s co-founder & research director Ilya Sutskever.

Crazy&Weird:

Backstory: I wondered if people would like to see something ‘crazy&weird’ in this newsletter and the votes told me ‘Yes’. So here we go:
AI & THE FUTURE OF MANUFACTURING.

[2025: A factory in China. There are no lights and thousands of industrial robots work in a complex, symphony. All you hear is the steady fizz of robotic movement.]

Machine view: multiple reinforcement-learning agents run simulations of the line in a large data center attached to the factory. They explore multiple perturbations of the manufacturing process, endlessly simulating the workload. When they discover more efficient approaches they initiate a High Priority Resource Call to the hardware scheduler in the data center and are assigned a chunk of computing resources to attempt to transfer their knowledge from the simulation into the real robots on the line. After the transfer is complete the robotic line reconfigures itself to account for the new simulation. Any errors are spotted by a thousand cameras staring down at the line. If the AI can diagnose the error it re-runs the simulation and comes up with a fix. If it can’t it sends the images&data of the flaw out to a large Mechnical Turk marketplace where human engineers observe the fault, come up with a fix, and send it back to the line. The system re-optimizes. Meanwhile, one line of the factory attempts to come up with perturbations of the assembled product, inventing wholly new versions of the devices by navigating through the latent feature space of the products. When new ‘Candidate Products’ are found it runs it through a series of tuned, expert systems and, if it gets a high enough score, simulates the product in a high-fidelity simulation. If it passes those tests then a Candidate Product is produced, airlifted by drone to a nearby human focus group and, if it satisfies their criteria, is sold on an EBay-like auction site frequented by the factory’s thousands of distributors. A bidding process takes place and in a few days/weeks data comes back about how the product succeeds in the market. If it does better than the existing product more parts of the factory are dedicated to creating new products in this style and the improvisation line begins exploring the latent space of the new products. In this way the manufacturing process begins to evolve according to a Cambrian evolution process, with AI automating much of the product R&D process.

Import AI: Issue 16: Changes at Twitter Cortex, Catastrophic Forgetting, and a $1000 bet

by Jack Clark

AI Term of the week: Catastrophic Forgetting: when neural nets completely and abruptly forget previously learned information upon learning new information

Reusability: One of the reasons why AI progress is accelerating is the community is creating more and more reusable components that can be plugged into different domains, frequently attaining performance equivalent to or better than hand-designed algorithms. The 2015 ImageNet challenge was won by a Microsoft system built out of Residual Networks, then in 2016 Microsoft made a speech processing breakthrough via a system that also relied on Residual Networks. Similarly, DeepMind’s WaveNet system has been slightly tweaked and re-applied to the domain of neural machine translation (PDF). This kind of re-use is a good thing as it suggests we are beginning to create the right sorts of low-level primitives that general intelligences can be built out of.

Domain conversion: We’re also seeing researchers work to convert hard problems into domains where we can use neural networks to learn to solve aspects of the problems. This brings more problems into a scope where we can train computers to learn how to solve problems that we’d have trouble programming by hand. For example, a recent OpenAI paper called Third Person Imitation Learning (PDF) is able to convert the problem of copying the actions of another entity into one amenable to generative adversarial networks, which let you learn a mapping between the behaviors of a teacher and a student without specific labels or additional data.

The Rude Goldberg AI Factory: startup  Tend.ai claims to have developed software lets you train robots to operate your 3D printers and associated genius-in-garage gear. The company says its software works with any robot, any webcam and any gripper, to imbue machines with the smarts to read the display and press the buttons on common factory items, like 3D printers. This could lead to a world where the mythological internet startup in a garage becomes the mythological internet-startup-slash-steampunk-robot-manufactory in a garage.

Who wants to make $1000? Venture capitalist Keith Rabois says AI is getting so good that it’ll soon be able to write screenplays and articles that are indistinguishable from those produced by humans. This is unlikely to happen in the short or medium-term — language remains one of the great unsolved challenges in AI. Designing good language systems requires an AI that can tie concepts in language to things it knows about the world, which requires an immense amount of what experts call ‘grounding’. We aren’t there yet. So AI researchers may want to take Keith Rabois up on his bet with Salesforce/Metamind’s Stephen Merity that he can “send you a manuscript that passes a Turing test.” Merity is in, as is Quora’s Xavier Amatriain. Anyone else?

Changes at Twitter: Twitter’s AI strategy is changing, judging by the recent departures of Hugo Larochelle, Ryan Adams, and others. The company had been trying to build an academic research group similar to those at Google, Facebook, and Microsoft. With their departures it seems the emphasis is now firmly on applied AI.

Gender bias in astronomy: researchers from the Swiss Federal Institute of Technology in Zurich use machine learning techniques to analyze 200,000 papers from 1950 to 2015. The key result: a paper whose first author is a woman is likely to get about 10 percent fewer citations than those that have men as the first authors. On the plus side, the fraction of papers where the first author is female has grown from less than 5% in the 1960s to about 25% today.

A vast machine intelligence landscape: Bloomberg Beta has published its third landscape of machine intelligence. The chart includes a third more companies than a year before and depicts a teeming menagerie of companies and organizations striving to apply ML to every component of the technology stack. “It feels even more futile to try to be comprehensive, since this just scratches the surface of all the activity out there,” they write.

Salesforce starts publishing AI papers: It’s normal for the research labs operated by the consumer companies to publish Ai research papers. Now that same spirit of openness seems to be moving into the enterprise as well. Recent Salesforce acquisition Metamind has published a number of papers recently on areas like question answering, natural language processing, recurrent neural networks, and more. It has also topped the leaderboard on the Stanford Question Answering Dataset (SQuAD) competition.

OpenAI bits&pieces:

How open should companies like Google be with their AI systems, what kinds of monopolies could we see emerge as a consequence of the AI revolution, and how do these issues relate to Japanese cucumbers? Those are some of the things Tim Hwang of Google and myself for OpenAI discussed in Toronto last month, which you can see in this video.

Import AI: Issue 15: Machine learning karaoke, Musk’s call for Basic Income, AI VS CRISPR

by Jack Clark

Do Androids go to the karaoke bar? Computers are getting much better at generating music and melodies thanks to the use of neural networks. New research from the University of Toronto, ‘Song From PI: A Musically Plausible Network for Pop Music Generation’ sees AI create some quite convincing songs. Listen to some samples here; they sound like the Muzak a sentient elevator might hum to itself. It uses a hierarchical recurrent neural network to develop songs with long-range structure; capturing the sorts of melodies and key changes that exist over multiple bars has been challenging to AI in the past, so this work is a step in the right direction. The authors also pair the tech with text/image synthesis AI to give the robots lyrics formed of the descriptions of images. “We were barely able to catch the breeze at the beach and it felt as if someone stepped out of my mind,” the robots sing. Just wait till we combined the synth-voice tech with Wavenet!

Automatic compression: the field of data compression is beginning to be revolutionized by AI, as new techniques like recurrent neural networks and autoencoders make it possible to train systems to compress and decompress images. New research (PDF) by Twitter Cortex takes a step in this direction, though there isn’t yet truly convincing performance relative to traditional compression methods.

People will wear masks when they speak of revolutionary things: spare a thought for spies, who will no longer able to meet in a loud bar to whisper clandestine truths to each other across a table. New research by the University of Oxford and DeepMind has created AI software (called LipNet (PDF)) that can learn to read people’s lips with an accuracy of around 93%, outperforming human experts. Important caveat: the training data consists of simple sentences with a limited vocabulary, such as ‘place blue in m 1 soon’, so it doesn’t work on real-world data yet. Check out this video to get an idea of how it works. “It’s a limited vocabulary & grammar dataset developed by colleagues at @shefcompsci, I know cos I’m in it,” writes Sheffield/Amazon’s Neil Lawrence.”So while the model may be able to read my lips better than a human, it can only do so when I say a meaningless list of words from a highly constrained vocabulary in a specific order.” (The technical term for this is ‘gibberish’). To take this sort of technique to the real world researchers will need to do three things: gather a vast amount of real-world video, develop software to be able to read lips from multiple angles and in varying lighting conditions, and build a language model that is sophisticated enough to be able to guess at the kinds of phrases people are using. The technology has such obvious utility, though, that it seems inevitable to be built. At some point in the future if you want to tell people secrets you’ll have to wear a face-mask, not because you’re afraid of pollution, but because you fear the decoding-ability of random CCTV cameras and airborne spy drones.

AI will deliver on the tech boom’s promise: “Economists have long wondered why the so-called computing revolution has failed to deliver productivity gains. Machine intelligence will finally realize computing’s promise,” argue James Cham and Shivon Zillis of Bloomberg Beta. That’s because ML will augment any business that involves a) computation and b) the gathering of data. Any successful business has this trait to one extent or another, and some of the lagrest and most significant enterprises have structured their companies to optimize for a) and b).

Battle of the world-changing technologies: AI VS CRISPR: It’s Election Week in America, so in that spirit I decided to conduct a terrifically biased and flawed poll. I asked people to tell me which of AI, CRISPR – a powerful gene editing technology -, or “something else” would have the biggest effect on the world over the next 10 years. More than 50 percent of the 700 votes (of my heavily AI-biased followers) went to AI, with the rest split between CRISPR and something else. A variant of my poll which added the option of ‘climate change’ saw votes split more evenly between AI and Climate Change, followed by CRISPR. We’er lucky to live in a time where we have not one but two totally transformative technologies developing at a rapid pace. It seems likely that CRISPR and AI will overlap over the coming years, as scientists use AI to better analyze the results from CRISPR experiments, and CRISPR lets us enhance our understanding of genes to the point where we can start turning them into organic computing machinery. The convergence of these fields will pave the way for things like ‘Turing biocircuits – programmable biological network pathways for partitioning and cycling energy and matter’. Turing biocircuits! What do you think?

The future is… militarized AI: the US military sees the deployment of AI & robotics as being as strategically important as its earlier development of nuclear weapons and precision munitions, according to the the Financial Times. (I’ve also heard from people in the defense universe that cognitive enhancements – whether that be drugs, or some kind of neural lace – are being viewed as key technologies for a future strategic advantage.)

Computers open their eyes: Advances in AI mean ‘images will be as transparent to computers as text is today’, says Benedict Evans of A16Z. What does this new world look like? Work by Mario Klingemann for the Google Arts Project gives us a peek. The XDegreesofseparation project uses software to build a visual bridge between radically different art objects, selecting a set of several pieces of art that help you go from the baleful geometry of an Aztec death mask to the sweeping, fluid curves of the Venus de Milo, or from an old painting of a woman to an ancient sculpture. The technical term for what’s going on here is ‘interpolation’ – the AI has learned to look at visual traits of an object and navigate between them. Expect more work like this, especially as research into a field known as generative models lets us navigate more of the mysterious landscape that unites all aesthetic objects.

AI develops a conscience: the development of AI has outpaced the development of ethics relating to AI. That’s created unfortunate situations where algorithms have displayed gross biases like echoing racial stereotypes in text-association programs or displaying rank sexism in language engines. But all is not lost; the community has woken up to these problems and researchers across the industry are devoting time to fixing things. Now they’re getting outside help as well: CMU has been given $10 million by a law firm called K&L Gates to set up a center to study ethics and AI. This center will sit  alongside related research efforts at Berkeley, Oxford, Cambridge, Stanford, and others to look at the societal and ethical impact of AI. In related news, the University of Oxford is hiring a researcher to look into the issues of personal safety, ethics, and security with regards to the internet of things, and how they relate to the amalgamation of a vast sea of digital data.

Moore’s Law is Dead, so now the chip-design rebels will stage their coup: When a great whale dies its carcass becomes an oasis in the belly of the ocean; for a while life flourishes amid its bones. The same is true of the end of different eras in technology. Now is the time of the death of Moore’s Law — and this is a good thing. Moore’s Law states that people can stuff double the amount of transistors in the same area of a chip every two years. That amazing law has driven much of the tech expansion in recent times, making intractable AI processing problems tractable, and putting supercomputers into the pockets of everyone with a smartphone. Now Moore’s Law is dying as ever-more-finely-detailed chip-making processes run into the uncaring and unyielding laws of physics. This means people are now exploring non-standard architectures for their processors, hoping to eke out performance through specialization rather than by relying on Moore’s Law. As one chip design student at MIT told me ‘it’s our time, now!’. That attitude drove Google to develop the TensorFlow Processing Unit chip for its data centers, led Microsoft to add FPGA co-processors to its Azure cloud, and motivated Intel to buy FPGA company Altera and AI-chip company Nervana. Now there’s a new chip startup called Graphcore which decloaked recently with $30 million in funding and plans to spread its Intelligence Processing Unit (IPU) chip and associated software far and wide. “The IPU has been optimized to work efficiently on the extremely complex high-dimensional models that machine intelligence requires. It emphasizes massively parallel, low-precision floating-point compute and provides much higher compute density than other solutions,” Graphcore writes.

The Kurt Vonnegut/Hunter S Thompson tech write you didn’t know you needed: This is your semi-yearly reminder to read James Mickens, the reassuringly insane computer science researcher and writer. “If a misaligned memory access is like a criminal burning down your house in a fail-stop manner, an impossibly large buffer error is like a criminal who breaks into your house, sprinkles sand atop random bedsheets and toothbrushes, and then waits for you to slowly discover that your world has been tainted by madness,” he says (PDF).  Read more of his work here. Thank me later.

OpenAI bits&pieces:

Automation requires a new social safety net: OpenAI co-chairman Elon Musk says the rise of automation and AI demands a new social safety net, specifically some kind of Universal Basic Income. There’s a vigorous debate going on among economists at the moment about whether UBI is viable, but one thing everyone agrees on is that advances in Ai are forcing a re-evaluation of what it means to work, what it means to be compensated for work, and how the government should react to the ensuing rise in inequality and centralizing of profits among a few operators of smart AI-infused capital. “Globalization may have ravaged blue-collar America, but artificial intelligence could cut through the white-collar professions in much the same way,” writes Jim Yardley in the New York Times.

ICLR-palooza: like many of our peers across the world we submitted a spread of papers to the ICLR conference. You can have a browse of them here. I’ll have a thorough writeup next week but the gist of this batch of research is about transfer learning for robots, more efficient reinforcement learning algorithms and further work on generative models.

Thanks for reading. If you have suggestions, comments or other thoughts you can reach me at jack@jack-clark.net or tweet at me@jackclarksf

Import AI: Issue 14: A Chinese robot boom, 1000-miles of driving data, and a culture clash at SoftBank

by Jack Clark

China’s 13th Five-Year Economic Plan promises big things for robots & AI: China plans to make substantial investments into robotics, according to the country’s 13th five year economic plan. The funding boost will support the Made in China 2025 initiative, which aims to improve the quality of Chinese-made goods while suffusing the country’s factories with smart machines. The robotics push may have significant consequences; in the 12th economic plan they said the development of a domestic semiconductor industry was a priority. Now, the majority of chips in the world’s fastest supercomputer use Chinese-designed ‘Sunway TaihuLight’ chips, rather than Intel’s which are used in the vast majority of other top machines. If the various robotic investments triggered by the plan pay off then there could be big consequences for artificial intelligence. Modern AI techniques have already been used by robot makers like Fanuc to substantially increase the speed at which a factory robot can be taught to excel at a particular task. This has led to related investments in areas like reinforcement learning (see: Fanuc’s partnership with Japanese AI startup Preferred Networks, or ABB’s relationship with Vicarious), which has fed into mainstream AI development. A flood of money from China, paired with the built-in end-customers in the form of large Chinese government labs, means robot-related AI development could speed up. Take a look at this gallery of a robot expo hosted by the Chinese army and imagine what kinds of advances will be possible by piggybacking on the manufacturing robot push.

If ‘Pepper’: then AI  == False: SoftBank’s tablet-clutching, Michelin Man-alike robot Pepper has not done nearly as well as the company hoped, disappointing customers. AI (or the lack of it) is partly to blame. The company had planned to apply modern AI techniques using neural networks to create a more advanced brain for the robot, but slowdowns and culture clashes between Aldebaran Robotics, a French company bought by Softbank in 2012, and the other engineering teams and managers (Japan), led to problems and delays. For example, instead of using deep learning techniques to learn to recognize emotions and react appropriately, the programmers had to dictate the responses by writing a laborious set of if-then statements. This limited Pepper’s abilities, and many companies wound up using the robot more as a cute tablet-toting marionette than as an interactive AI assistant. It’s a lesson in the importance of communication and what happens if that goes awry.

More gratuitous GTAV self-driving bot footage: In Issue 12 I wrote about research that shows hit game GTAV is a reasonable simulator to train self-driving cars in. This week I discovered that Craig Quiter, a technical contractor at OpenAI, has a personal project called DeepDrive where he uses GTAV to develop his own car systems. Watch videos of a car navigating traffic, keeping to the center of a lane, and marvel at the moment a snowstorm causes its visual brain to break down. Videos & more information here.

I would drive 500 miles, and I would self-drive 500 more, just to be the (self-driving) car that brings data to your door: the proliferation of free datasets for self-driving continues. The Oxford Robot Car dataset comes from researchers spending a year traversing the city of Oxford twice a week in a self-driving car. The 20 terabyte release involves around 100 journeys on the same route and covers both camera data, as well as LIDAR and GPS. The repetition of the route is important because the resulting dataset will let people teach AI systems how to make sense of the many permutations of route world, like shifting pedestrians, different weather patterns, rain, and so on. This follows similar data releases from Udacity and Comma.ai. We’re lucky to live in a time where companies and universities generate so much free data.

Comma catastrophe self-driving car startup Comma.ai has cancelled its main product, the Comma 1 – a $999 kit to retrofit modern cars into self-driving robo-chauffers. Founder George Hotz said he cut the product after receiving a letter from US regulator NHTSA. “comma.ai will be exploring other products and markets. Hello from Shenzhen, China. -GH”, Hotz writes.

Building up Canada’s AI ecosystem: In the same way PayPal spawned a generation of entrepreneurs who came to influence the tech industry, Canada created its own gang of hugely influential AI researchers through CIFAR, a funding body that supported people like Geoff Hinton (now at Google), Yann Lecun (now at Facebook), and Yoshua Bengio (now at UMontreal). Their contributions form much of the bedrock of the current popular AI techniques, ranging from systems for image recognition and machine translation at Google, to state-of-the-art memory systems at Facebook. But Canada has struggled to retain its top talent, as professors are drawn to work at other companies, and US-based schools massively increase investment in AI programs. To counter that, a group of researchers and entrepreneurs have launched Element AI, a Montreal-based AI incubator slash research lab. Influential research Yoshua Bengio, the Paul Erdös of AI, is one of the founders.

Welcome to the world of the self-defending autonomous AI corporation: in Charles Stross’ (excellent, free) sci-fi novel, Accelerando, the world is suffused with smart, semi-autonomous AI-driven digital corporations that indulge in a constant cacophony of deal-making, legal subterfuge, and company formation and destruction. It feels like a plausible future, but requires two prerequisites: 1) a form of digital currency with various forms of metadata built into it to let computers participate in a universal economic market, 2) the ability of corporations to exchange information with each other privately and efficiently. 3) better decision-making AI systems capable of feats of memory, transfer learning, and ideation. Software like Bitcoin and Ethereum attempts to solve problem one, new research from Google tackles option two, and the AI research community is working on option three. The new Google paper, Learning to Protect Communications with Adversarial Neural Cryptography, outlines what they call a neural cryptography system. This lets them have semi-independent AIs improvise a secure communication channel with each other in the presence of an adversary. Give it a few years and they’ll be a proliferation of microscopic economic agents that barter privately with each other in a market of digital information.

Image recognition isn’t only for the data titans: Image recognition startup Clarifai has raised a $30 million Series B. Clarifai is led by AI researcher Matt Zeiler, and the company proved its AI chops in 2013 by winning the ImageNet competition. It sells image recognition services to people around the world, letting people use AI to automatically organize their photos and videos, or identify specific items in images (hypothetical example: ad agencies training a Clarifai AI to recognize and spot different brand logos on clothing, then using the software to automatically patrol the web to identify photos containing the brand. “Today Clarifai not only hosts a static API that tags thousands of images and videos with human level accuracy, but now empowers anyone from a wedding photographer to a large retail company to a sail boat enthusiast (hat tip: USV’s Albert Wenger) to easily access and build products employing Google level AI with drag and drop ease of use,” writes Lux Capital, which invested in the round. Clarifai’s ongoing success is an intriguing counterpoint to the narrative that it’s difficult for AI startups to compete with the resources wielded by vast tech companies like Amazon, Google, Microsoft, and others.

Import AI: Issue 13: Microsoft’s speech breakthrough, how to make neural nets that resist interrogation, and AI-generated Halloween art

by Jack Clark

Teaching old neural nets new tricks: Congratulations to Geoff Hinton, a self-described “machine learning fossil” who, along with collaborators, has fleshed out an idea he has been working on since 1973. The research, ‘using fast weights to attend to the recent past’ tries to make neural networks a little bit more brain-like through the use of a ‘fast weight’. “These “fast weights” can be used to store temporary memories of the recent past and they provide a neurally plausible way of implementing the type of attention to the past that has recently proved very helpful in sequence-to-sequence models,” according to the  research paper’s abstract.There’s a good, thorough lecture of the approach by Geoff Hinton here. Along with being quite smart Hinton is also reasonably funny and, had he not squandered his life on AI, could have been a very good stand-up comic. This represents yet another name on the lineup of the current AI-Memorypalooza, and will be appearing besides Memory networks, differentiable neural computers, LSTMs, and other memory-oriented systems in a workshop near you soon.

My fair Microsoft: Microsoft claims to have exceeded human performance at recognizing speech. This has taken a few experts by surprise. “I must confess that I never thought I would see this day,” says British-American linguist Geoff Pullum. The system beat human experts at identifying sounds on Switchboard, a long-in-the-tooth audio dataset consisting of over 2,400 telephone conversations among 543 people from the United States. There’s evidence that this data doesn’t capture all the nuances and difficulties of speech, though it does contain salt-of-the-earth phrases like ‘it really chips really easy” and “adept at doing things, it’s just…”. The true test, though, is on real-world datasets. Chinese search giant Baidu, for instance, gathered more than eight thousand hours of audio data to test and train its Deep Speech system, and Google harvests vast amounts of audio information to build its own classifiers as well. Perhaps we need a new ASR dataset?

Machine learning & fraud: one of the areas where machine learning is going to be widely deployed is in fraud identification. Fraud is almost the perfect problem for ML approaches because it involves visualizing odd permutations in a vast stream of data. So it’s unsurprising to see the launch of Stripe Radar. The bigger implication of services like this is that it gets more effective as more data gets plugged into it, so as a customer grows the predictive capabilities of Radar will grow as well, and there’s also a good chance that the entire platform will benefit from insights gleaned from each individual customer. This is why it currently looks like ML will turbocharge the benefits that companies extract from operating widely used platforms.

Open research: Francois Chollet, the creator of Keras, appears to really, really enjoy working every hour in the day, given that his new side project is the ‘Artificial Intelligence Open Network’. AI ON lists open problems in AI and gives people the opportunity to work on real, meaningful problems with senior AI researchers. It’s somewhat like OpenAI’s Requests for Research, though has a more well-defined open research process. This fits with the general tendency towards openness and transparency in AI. Long may it continue!

AI & the re-evaluation of intellectual property: spare a tear for the lawyers who will soon grapple with the myriad intellectual property issues brought about by the rise of AI. “How will we protect the intellectual property embodied in those products and services, if anyone can reverse engineer their core IP simply by using them and feeding their output into commodity machine learning systems”? writes Daniel Tunkelang. The answer could be to design systems that are resilient to such model-scraping attacks — there’s already research being done here, including a contribution by OpenAI (more on that at the bottom of this letter).

Brain augmentation is closer than you thinksays Braintree founder Bryan Johnson who is pouring $100 million of his money into Kernel, a company that aims to create ‘the world’s first neural prosthetic for human intelligence enhancement’. Between that and recent work on ‘neural lace’ technology it seems that the era of brain-fiddling is upon us. Let’s hope we can avoid the situation rendered in Black Mirror S3:E2 ‘Playtest’.

Mo’ AI, Mo’ Macroeconomic Demand-Side Problems: “A significant number of tasks now performed by humans will be performed by machines and artificial intelligence. We could very well see 5 million jobs eliminated by the end of the decade because of technology,” says Andy Stern, former president of the Service Employees International Union (SEIU).

Normalization-on-normalization-on-normalization: First there was batch normalization, then layer normalization, and now there is streaming normalization (PDF). These techniques make it more computationally efficient to train neural networks. Though batch normalization was only outlined in 2015 it has already become quite widely used. We may expect the same for layer and streaming normalization. “Machine learning systems often get confused by stimuli that vary wildly but in uninteresting ways, for example, a visual recognition system might get confused by a room where the lights get randomly brighter and dimmer but nothing else about the room changes. Batch normalization is a technique that helps us to ignore such uninteresting variation, both at the level of the network’s raw inputs, and its higher level judgements,” says OpenAI’s Dario Amodei.

Chinese cash for UK AI talent: A Chinese private equity firm will partner with UK-based Founders Factory to invest in UK AI startups, giving China access to some of the UK’s excellent AI talent, and the UK companies access to China. In Beijing, entire city blocks have been converted from focusing on cloud computing to instead focus on AI research, according to Miles Brundage.

Halloween AI: Welcome to MIT’s ‘Nightmare Machine”, some dastardly AI-fiddlers have taken it upon themselves to use style transfer techniques – a way of getting a neural network to interpret the aesthetic style of one picture and apply it to another one – to create pictures of ‘haunted faces’ and ‘haunted places’. Take a tour of the ghoulish AI creations, but remember ‘images on this website are generated by deep learning algorithms and may not be suitable for all users. They contain scary content’. Scary content! If you’ve always thought the poster for the film Jaws could be improved by the addition of loads of skulls then this will be right up your alley. Perhaps we’ll soon see a similarly spooky filter arrive in popular app Prisma?

/// OpenAI bits&pieces ///

Developing machine learning systems that can ingest sensitive data and output anonymized answers is a challenge. In an ideal world, you want to limit the distribution of the personal data – say, someone’s medical information – so that people can’t try to exploit flaws in your ML model to uncover private information. That was the motivation behind research from Penn State’s Nicolas Papernot,  OpenAI’s Ian Goodfellow, and several people from Google. The research relies on a system where multiple ‘teacher’ networks are trained on a dataset, then give slightly garbled predictions to a ‘student’ network, which learns to classify things without having ingested any identifiable data, making it resilient to model-extraction attacks. “The approach combines, in a black-box fashion, multiple models trained with disjoint datasets, such as records from different subsets of users. Because they rely directly on sensitive data, these models are not published, but instead used as teachers for a student model. The student learns to predict an output chosen by noisy voting among all of the teachers, and cannot directly access an individual teacher or the underlying data or parameters,” they explain in the paper: ‘Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data’.

Reinforcement learning presentation: OpenAI researcher John Schulman gave a presentation at Galvanize in SF this week on deep reinforcement learning through policy optimization. Slides of the (non-recorded) talk are available here.

Korean robots in OpenAI Gym: A Korean organization is training a robot to walk using our open source OpenAI Gym software. Facebook’s translation isn’t quite up to the task of providing a non-garbled translation, though…

/// Administrative Note:Two weeks ago I asked for advice about what people would like to see from an OpenAI newsletter. I got lots of helpful responses and am now vigorously digesting them. Thank you! ///

Import AI: Issue 12: Learning to drive in GTAV, machine-generated TV, and a t-SNE explainer

by Jack Clark

Q: ‘Why did your self-driving car just go through a red light?’ A: ‘Because it was trained in Grand Theft Auto 5, officer.’ Some folk wisdom about AI research is that only a few companies have the wherewithal to build up the tremendous stores of information necessary to develop world-changing AI systems, like self-driving cars. This was true for a while but is now changing as researchers get better at teaching computers to use data from simulated environments. A new paper, called ‘Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks’ (PDF), describes a way to generate training data for a self-driving car by pulling screenshots & depth information & object labels from popular videogame Grand Theft Auto 5. The researchers harvest data from the game to create a self-driving car vision model that performs comparably to ones made of real-world data. They’re even able to tap into GTAV’s complex weather systems to get driving data in a variety of conditions, like fog, rain, haze, snow, and so on. This research implies that the competitive walls around self-driving car development are somewhat lower than people realize. The same phenomenon is taking place in robotics with new papers from DeepMind (PDF) and OpenAI (PDF) outlining ways to train AI systems in simulators then transfer into the real world.

Demystifying t-SNE: t-SNE is a powerful tool for visualizing the sorts of high-dimensional datasets, but developing good intuitions about what it shows you is extremely difficult. That’s mostly because we’re trapped in three spatial dimensions and so trying to develop a mental model of a 500-dimensional data representation is hard, like an ant having to grok the concept of high-frequency trading. This thorough explainer may prove helpful.

300 lines of code can build an AI agent that learns use a deep deterministic policy gradient algorithm to drive a car.

The future is… a million machine-generated episodes of Cheers, Friends, and Happy Days: research from the University of Leeds shows how to build ‘virtual talking avatars of characters fully automatically from TV shows’ (PDF).The approach use a generative model to sample the style of speech and video appearance of a character from a TV show, letting you, say, re-animate Joey from Friends or Norm from Cheers and make them do and say unspeakable things. Once the rest of the AI community develops better language models this could be used to create endless, machine-generated television shows. It could also work with current AI techs, though the results would be a little unsatisfying. As RNN-generated Joey says: ‘Seriously give me a clown on the table that’s all!’ *cue theme tune*

Memory & cognition: DeepMind has published a new Nature paper on its system that pairs differentiable memory with neural networks. The approach lets computers teach themselves about interconnected concepts, like nodes in a transit system or a family tree, and then reason about them. The differentiable neural computer (PDF), is based on earlier work called the Neural Turing Machine. (Due to the immense lag imposed by traditional, paywalled publishing this paper is relatively old, having been submitted in January of this year. Since then we’ve seen new memory-based techniques from DeepMind, the University of Montreal, Facebook, and others.) Additional perspective on DeepMind’s tech from Facebook AI scientist Yuandong Tian here. Tian also reveals that Facebook will shortly publish a paper about its DOOMBOT.

You have information, consider donating it:: Do you work in AI? If so, please consider filling out this survey. One of the things the AI community lacks is good information about its progress, habits, composition, and methods. The more information we can share, the faster we can progress!

Transcontinental data vein: from the perspective of an alien the story of the internet is really the story of the world being connected by an ever-growing tangle of data-conducting cables, propagating information across an otherwise fractured planet. So they’d probably cheer at seeing Facebook, Google, Pacific Light Data Communication, and TE SubCom, team up to build a 12,800km-long 120Tbps cable between Hong Kong and Los Angeles. If that piques your interest then you might like this photo tour aboard an Alcatel-Lucent cable ship.

Machine learning is the new statistics because it helps you achieve a more sophisticated REALITY ANALYSIS LEVEL.

Better drug design with machine learning: One of the dreams of AI researchers is to invent software that can conduct scientific experiments (whether this is due to laziness or curiosity on their part is less clear.) A new paper called ‘Automatic chemical design using a data-driven continuous representation of molecules’ (PDF) goes down this path. The system uses similar technologies used by Google to read your email and offer AI-generated responses (Smart Reply),to let scientists explore molecules and search for new drugs. The software converts molecules into a form (a fixed-dimensional vector) that machine learning approaches can deal with, making the immense combinatorial space of chemistry navigable to ML software. Once you have a vector representation of a molecule you can start to fiddle with the dials that describe its characteristics and use this to explore nearby molecules you may not have previously studied (the same principle works for words in translation systems, where you can fiddle with the dials of the representation of ‘cat’ and navigate to nearby entities like dogs, mice, and so on). The scientists are able to use this approach to discover some molecules that have even better properties than known ones. This brings us closer to an era where machines can help us to discover new drugs and treatments. Fascinating paper worth multiple reads.

The future is here and it is made of advertising. : Earlier this summer Uber advertised its services in Mexico City via drones that hovered above traffic, holding signs that scolded drivers. The future is here and it’s made of advertising.

/// OpenAI bits&pieces ///

Self-organizing conferences are surprisingly viable: OpenAI held its first self-organizing conference on machine learning and things went well. But don’t take my word for it, read this post from Victoria Krakovna about some of the AI&Safety things we talked about. You can find minutes from some of the other sessions on the wiki.

Language matters: Language is inextricably tied to the environment the thinking entity grows up in, so researchers are starting to design worlds that will encourage baby AI minds to develop their own language systems and in doing so (hopefully) become smarter. Facebook has published software tools and papers in this area, and both DeepMind and the University of Oxford have made great strides in having multiple agents learn to communicate with one another to solve problems. OpenAI is conducting research in this area as well; Jon Gauthier and Igor Mordatch have proposed (PDF) a way you might want to do this.

Matrix robots: A new paper, Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model (PDF) outlines a technique to train a robot in a simulator then transfer some of those insights into a real world machine, and lists some future research directions. We’re a few years off from ‘I know kung-fu’, but we’ll get there eventually.

Import AI: Issue 11: Robots learn through fighting&friendship, gloopy DNA storage, and an AI acronym crime

by Jack Clark

AI – hyped, or underhyped? AI is receiving an extraordinary amount of attention these days. Is this justified? Talk to technologists and the answer is ‘yes, with some qualifications’. People tend to feel like the press coverage of AI glosses over the numerous flaws, dead-ends, and implementation costs of modern technology. But at the same time most people are convinced that the commercial potential of AI is vast and mostly unexplored. That’s why Jeremy Howard, CEO of fast.ai, says in this video interview that ‘the potential for deep learning is greater than the potential of the internet in the early 90s’.

What is yellow, expensive, and marginally smarter than a rock? Fanuc’s robots! The company has begun to apply AI techniques like reinforcement learning to its robots so that they can be taught to do industrial tasks faster. Now Fanuc and NVidia have announced plans to stuff more GPU-based computing power into the yellow machines.

1 + 1 = SWARM INTELLIGENCE: Rapyuta Robotics, a spin-off from ETH Zurich, recently got $10 million in Series A funding to help it commercialize technology to let different robots learn from each other. Seems like the right time to do so, given this research from Google which shows how you can train multiple robots to solve the same task and in doing so learn more efficiently than if you were just training on a single machine. But collaboration might be altogether too boring. Just wait until the robots start to teach each other to solve tasks by attempting to outfox one another, as outlined in this tantalizing paper from CMU and Google. ‘Having robots in adversarial setting might be a better learning strategy as compared to having collaborative multiple robots,’ they write in Supervision via Competition: Robot Adversaries for Learning Tasks.

Samsung buys Viv: Samsung has acquired Viv, an AI startup founded by the people who helped create Siri at Apple and before that worked on SRI’s Calo project. Viv generated a lot of press and made frequent cryptic references to work done in program synthesis but did not publish any meaningful technical details about its approach. Now that is has been acquired I hope the company could publish a paper so the AI community can assess its work and share in any insights the team has had.

No AI’s in the classroom, or else! As if phones weren’t bad enough the Allen Institute for AI has released a live demo of Euclid, a tool to solve SAT-style math questions. It’s got some weird tendencies, for example: ‘Question: What is the smallest number? Euclid: -120.0’. Well, that settles that then…

Even more free data: Self-driving startup Comma.ai released 80GB of driving data a couple of months ago. ‘Pah! That’s nothing,’ I imagine Udacity’s Oliver Cameron saying, as he presses the big red button to release 223GB of Mountain View driving data. Sooner or later we’re going to have trouble storing all of this information, so keep an eye on the burgeoning field of DNA storage for future solutions to density, redundancy, and resiliency problems. “We stored an entire computer operating system, a movie, a gift card, and other computer files with a total of 2.14*10^6 bytes in DNA oligos. We were able to fully retrieve the information without a single error even with a sequencing throughput on the scale of a single tile of an Illumina sequencing flow cell,” write some gloopy researchers in the abstract to their paper ‘Capacity-approaching DNA storage’.

Import AI + OpenAI: I’m planning a regular blogpost/newsletter for OpenAI in which I’ll try and analyze the monthly trends in AI from both a research and industry perspective. Is there anything in particular you think I should focus on? Get in touch, please! It’ll be a bit longer than an Import AI issue and a bit more technical, I think.

New research: GAWWN$%^@? Call the acronym police, a crime has been committed! A new paper proposes the Generative Adversarial What-Where Network (GAWWN). Like other GANs it can create synthetic images that seem plausible, and unlike other approaches it can follow detailed instructions, creating better, more realistic images than before. ‘This is pretty bonkers,’ says Miles Brundage. We agree!