{
  "SchemaVersion": 1,
  "IndexedAtUtc": "2026-09-17T00:09:07.2610287Z",
  "MatchingPolicy": "Only verified references are linked. Similarity, shared authors, topical overlap and year alone do not establish a citation. Publication versions remain distinct unless explicitly documented.",
  "Entries": [
    {
      "Slug": "mcculloch-pitts",
      "Paper": "A logical calculus of the ideas immanent in nervous activity",
      "AtlasYear": 1943,
      "Status": "indexed",
      "Method": "publisher-reference-section",
      "SourceUrl": "https://link.springer.com/article/10.1007/BF02478259",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "Publisher references",
          "StartLine": 1,
          "EndLine": 3,
          "PdfPages": [],
          "Text": "Carnap, R. 1938.The Logical Syntax of Language. New York: Harcourt, Brace and Company.\nHilbert, D., und Ackermann, W. 1927.Grundüge der Theoretischen Logik. Berlin: J. Springer.\nRussell, B., and Whitehead, A. N. 1925.Principa Mathematica. Cambridge: Cambridge University Press."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference section extracted from the linked source. Only reviewed matches to existing Atlas works become citation links; extraction may retain typographic or column-order artifacts."
      ]
    },
    {
      "Slug": "hebb",
      "Paper": "The Organization of Behavior: A Neuropsychological Theory",
      "AtlasYear": 1949,
      "Status": "indexed",
      "Method": "archive-ocr-and-windows-ocr",
      "SourceUrl": "https://archive.org/download/in.ernet.dli.2015.226341/2015.226341.The-Organization.pdf",
      "PdfSha256": "D6A00C6010AAB2D905FB6DAB3EADFAE10A1318A5842E929CD9CFF068965B3C46",
      "Sections": [
        {
          "Section": "Bibliography, printed pages 305-319",
          "PdfPages": [
            328,
            329,
            330,
            331,
            332,
            333,
            334,
            335,
            336,
            337,
            338,
            339,
            340,
            341,
            342
          ],
          "Text": "Bibliography\n\n\nAdams, D. K. 1929. E.'cperimental studies of adaptive behavdor in cats.\nComp. Psychol. Monog., 6, ’No. 1.\n\nAdrian, E. D. 1931. Potential changes in the isolated nervous system of\nDytisciis marginalis. J. Physiol., 72, 132-151.\n\nAdrian, E. D. 1934. Electrical activity of tire nervous system. Arch.\nNeurol. Psychiat., 32, 1125-1136.\n\nAdrian. E. D , and Buytcndijk, F. J. J. 1931. Potential changes in tire iso-\nlated brain stem of the goldfish. J Physiol., 71, 121—135.\n\nAdrian, E. D., and Matthews, B. H. C. 1934. The interpretation of po-\ntential waves in the cortex. J. Physiol., 81, 440-471.\n\nAllen, C., and Broster, L. R. 1945. A further case of paranoid psychosis\nsuccessfully treated by adrenalectomy. Brit. Med. J., No. 4402, 696-698.\nAllport, G Vi^ 1946. Effect: a secondary principle of learning. Psychol.\nRev., 53, 335-347.\n\nAnderson, J. E. 1939. The limitations of infant and preschool tests in the\nmea.surement of intelligence. J. Psychol, 8, 351-379.\n\nArvanitaki, A. 1942. Effects evoked in an axon by the activity of a con-\ntiguous one. J. Neurophysiol., 5, 89—108.\n\nBard, P. 1934. On emotional expression after decortication with some re-\nmarks on certain theoretical views. Psychol Rev., 41, 309-339.\n\nBard, P. 1942. Neural mechanisms in emotronal and sexual behavior.\nPsychosom. Med., 4, 171-172.\n\nBartley, S. H., and Bishop, G. H. 1933. Factors determining the form of\nthe electrical response from the optic cortex of the rabbit. Amer. J.\nPhysiol, 103, 173-184.\n\nBartley, S. PL, and Chute, E. 1947. Fatigue and impairment in man.\nNew York; McGraw-PIill.\n\nBeach, F. A. 1937. The neural basis of innate behaviaa'. 1. Effects of\ncortical lesions upon the maternal behavior pattern in the rat. J. Comp.\nPsychol, 24, 393-439.\n\nBeach, F. A. 1939. The neural basis of innate behavior. III. Compari-\nson of learning ability and instinctive behavior in the rat. J. Comp.\nPsychol, 28, 225-262.\n\nBeach, F. A. 1942. Analysis of factors involved in the arousal, mainte-\nnance and manifestation of sexual excitement in male animals. Psy-\nchosom. Med., 4, 173-198.\n\nBeach, F. A. 1947a. A review of physiological and psychological studies\nof so.xual behavior in mammals. Physiol Rev,, 27, 240-307.\n\n305\n\n\n\n306 Bibliography\n\nBeach, F. A. 1947h. Evolutioiiaiy changes in tlie physiological control of\nmating behavior in mammals. Psychol. Reo,, 54, 297-315.\n\nBeach, F. A. 1948. Hormones and behavior. New York: Hoeber.\n\nBeliak, L., and Willson, E. 1947. On the etiology of dementia piaecox\n• • •. J. Neio. Meat, Dis., 105, 1-24.\n\nBellows, R. T. 1939. Time factors in water drinking in dogs. Amer. J.\nPhysiol, 125, 87-97.\n\nBirch, H. G. Ib/s. The relation of previous experience to insightful\nproblem-solving. J. Comp. Psychol, 38, 367-383.\n\nBishop, G. H. 1946. Neive and synaptic conduction. Ann. Rev. Physiol,\n8, 355-374.\n\nV. Bonin, G., Garol, H. W., and McCulloch, W. S. 1942. The functional\norganization of the occipital lobe. In Muver, H., Visual mechanisms.\nBiol Stjmp., 7, 165-192.\n\nBoring, E. G. 1916. Cutaneous sensation after nerve-division. Quart. I.\nExp. Physiol, 10, 1-95.\n\nBoring, E G. 1930. A new ambiguous figure. Amer. 1. Psychol, 42,\n444-445.\n\nBoling, E. G. 1933. The physical dimensions of consciousness. New\nYork; Century.\n\nBoring, E, G. 1946. Mind and mechanism. Amer. 1. Psychol, 59, 173-\n192.\n\nBousfield, W. A. 1935. Quantitative indices of the effects of fasting on\neating-beha-vior. 1. Genet Psychol, 46, 476-479.\n\nBowman, K. M. 1935. Psychoses with pernicious anemia. Amer, J.\nPsychiat., 92, 371-396.\n\nBowman, K. M. 1946. Modern concept of the neuroses. 1. Amer. Med.\nAssoc., 432, 555-557.\n\nBridgman, C. S., and Smith, K. U. 1945. Bilateral neural integiation in\nvisual perception after section of tlic corpus callosum. /. Comp. Neurol,\n83, 57-68.\n\nBronk, D. W. 1939. Synaptic mechanisms in sympathetic ganglia.\n1. Neurophysiol, 2, 380-401.\n\nBrown, Warner. 1932. Spatial integraBons in a human maze. Univ.\nCalif. Puhl. Psychol, 5, 123-134.\n\nBruetsch, W. L. ,.1947. Rheumatic biain disease: late sequel of rheumatic\nfever. 1. Amer, Med. Assoc., 134, 450-454.\n\nCarlson, A. J. 1916. The control of hunger in health and disease. Chi-\ncago; Univ. Chic. Press.\n\nCarmichael, L., Hogan, H. P., and Walter, A. A. 1932. An exiierimental\nstudy of the effect of language on tire reproduction of vi.siially perceived\nform. J. Exp. Psychol, 15, 73-86.\n\nChariuler, A. R. 1934. Beauty and human nature. New Yoik; Apijleton-\nCentuiy.\n\nClark, G., and Lashley, K. S. 1947, Visual disturbances following frontal\nablations in the monkey, Anat. Rec., 97, 326.\n\n\n\n307\n\n\nBibliography\n\nCobb, S. 1944. Personality as affected by lesions of the brain. In Hunt,\nJ. McV., Personality and the behavior disorders. New York; Ronald,\nVol. I, pp. 550-581.\n\nCobb, S., Cohen, M. E., and Badal, D. \\V. 1946. Capillaries of the nail\nfold in patients with neuiocirculalory astlienia (effort syndrome, anxiety\nneurosis). Arch. Neurol. Fsychiat , 56, 643-650.\n\nCohn, R. 1945. Electroencephalographic study of prefrontal lobotomy.\nArch. Neurol. Psychiat., S3, 283-288.\n\nCowles, J. T., and Nissen, H. VV. 1937. Reward-expectancy in delayed\nresponses of chimpanzees. J. Comp. Psychol., 24, 345-358.\n\nDandy, W. E. 1933. Physiologic studies following extirpation of tire right\ncerebral hemisphere in man. Bull. Johns Hopkins Hasp., S3, 31-51.\n\nDaniel, R. S., and Smrth, K. 15. 1947. The sea-approach behavior of the\n\nneonate loggerhead turtle. J. Comp. Physiol. Psychol,, 40, 413-420.\n\nDashiell, J. F. 1928. Are there any native emotions? Psychol. Rev., 35,\n319-327.\n\nDavison, C , and Dcmuth, E. L. 1945. Disturbances in sleep mechanism;\na clinico-pathologic study. II. Lesions at the corticodienceirhalic le\\'el.\nArch. Neurol. Psychiat., 54, 241-255.\n\nDonker, P G. 1946. Results of treatment of psychoneiiroses by the gen-\neral practitioner: a follow-up study of 500 cases. N. Y. State J. Med.,\n46, 2164-2166.\n\nDennis, W. 1934. Congenital cataract and unlearned behavior. J. Genet.\nPsychol., 44, 340-350.\n\nDennis, W. 1940. Infant reaction to restraint: an evaluation of Watson’s\ntheory. Trans. N. Y, Acad. Sci., Ser. 2., 2, No. 8, 202-218\n\nDenny-Brown, D. 1932. Theoretical deductions from the physiology of\nthe cerebral cortex. J. Neurol. Psychopathol, 13, 52-67.\n\nDewan, J. G., and Owen, T. 1945. Mental illness and the principles of\nmedicine. Canad. Med. Ass. J., 52, 349—357.\n\nDoll, E. A. 1933. Psychological signrlrcance of cerebral birth lesions.\nAmer. J. Psychol., 45, 444-452.\n\nDoll, E. A., Phelps, W. M., and Melcher, R. T. 1932. Mental deficiency\ndue to birth injunes. New York: Macmillan.\n\nDrew, G. C. 1938. The function of pum.shment in learning. J. Genet.\nPsychol., 52, 257—267.\n\nDubner, H. H., and Gerard, R. W. 1939. Factors conU'oUing brain po-\ntentials in the cat. J. Neurophysiol., 2, 142-152.\n\nDunlap, K. 1932. Habits: their making and remaking. New York:\nLiveright.\n\nEditors, Nutrition Reviews. 1944. Self-selection of diets. Nutrition Rev.,\n2, 199-203.\n\nEgaria, E., Johnson, R. E., Bloomfield, R., Brouha, L., et al. 1942. The\neffects of a diet deficient in the vitamin B complex on sedentary men.\nAmor. J. Physiol, 137, 731-741.\n\n\n\n808 Bibliography\n\nErlanger, J. 1939. The initiation of impulses in axons. ]. Neurophysiol.,\n2, 370-379.\n\nFeixaro, A., Arieti, S., and English, W. H. 1945. Cerebral changes in\nthe course of pernicious anemia and their relationship to psychic symp-\ntoms. J. Neuropathol. Exp. Neuiol., 4, 217-239.\n\nFields, P. E. 1932. Studies in concept formation. I. Comp. Psychol.\nMonog., 9, No. 2.\n\nForbes, A. 1939. Problems of synaptio function. J. Neurophi/siol., 2,\n465-472.\n\nFreeman. G. L. 1934. Introduction to plrysiological psychology. New\nYork. Ronald.\n\nFreeman, W., and Watts. J. W. 1942. Psychosurgery: intelligence, emo-\ntion and social behavior following prefrrfntal lobotormj for mental dis-\norders. Springfield: Thomas.\n\nFreeman, W., and Watts, J. W. 1946. Psychosurgery. In Spiegel, E. A.,\nProgress in neurology and psychiatry: an annual review. New York:\nGrune and Stratton, pp. 649-661.\n\nFuchs, W, 1920. Untersuchungen uber das Sehen der Hemianopiker und\nHemiamblyopiker: II. In Gelb, A., and Goldstein, K., Psychologtschen\nAnalysen hirnpatlwlogischer Fulle. Leipzig: Barth, pp. 419-561.\n\nFulton, J. F. 1943. Physiology of the nervous system. 2nd Ed., New\nYork: Oxfoid Univ. Press.\n\nGantt, W. H. 1938. Extension of a conflict based upon food to other\nphysiological systems and its reciprocal relations with sexual functions.\nAmer. J. Physiol, 123, 73-74.\n\nGantt, W. II. 1944. Experimental basis for neurotic belia\\'ior: origin and\ndevelopment of artificially irroduced distuibances of beliai'ior m dogs.\nPsychosioin. Mad. Monog., 3, Nos. 3 and 4.\n\nGarrison, M. 1947. The genetics of schizophrenia. J. Abn. Soc. Psychol,\n42, 122-124.\n\nGas.sei, H. S. 1937. The conb’ol of excitation in the nervous system.\nHarvey Lect., pp. 169-193.\n\nGellennan, L. W. 1933. Form discrimination in chimpanzees and two-\nyear-old cliildren: I. Form (triangularity) per se. I. Genet. Psychol,\n42, 3-27.\n\nGibbs, F. A, 1945. Electrical activity of the brain. Ann. Rev. Physiol,\n7, 427-454.\n\nGibson, J. J. 1929. The reproduction of visually perceh’ed forms.\nI. Exp. Psychol, 12, 1-39.\n\nGibson, J. J. 1941. A critical review of die concept of set in contempo-\nrary experimental psychology. Psychol. Bull, 38, 781-817.\n\nGibson, J. J., and Crooks, L. E. 1938. A theoretical field-analysis of\nauTomobile diiving. Amer. J. Psychol, SI, 453—471.\n\nGillespie, W. H. 1944. The psyclioneuroses. J. Menf. Set., .90, 287-306.\n\nGoodenough, F. L. 1931. Anger in young children. Minneapolis: Univ.\nMinnesota Press.\n\n\n\nBibliography 309\n\nGreene, R., Paterson, A. S., and POe, G. C S. 1945. Hypertrichosis with\nmental changes; the effect of adrenalectomy. Brit. Med. J., No. 4402,\n698-699.\n\nGuthrie, E. R. 1946. Psychological facts and psychological theory.\nP.st/clinl. Bull., 43, 1-20.\n\nHalstead, W. C. 1947. Brain and intelligence. Chicago: Unh'. Cliic.\nPress\n\nHanawalt, N. G. 1937. Memory trace for figures in ’recall and recogni-\ntion. Arch. Psychol , No. 216, 1-89.\n\nHams, H. J. 1944. Brucellosis: a case report illustrating a psychosomatic\nproblem. Psijchosom. Med., 6, 334—335. -\n\nHauplman, A. 1946 Cairillaries in the finger nail fold in patients with\nneurosis, epilepsy, and migraine. Arch. Neurol. Psychiat., 56, 631-642.\n\nHead, H. 1920. Studies in neurology. London: Frowde, Plodder and\nStoughton.\n\nHeath, R. G., and Pool, J. L. 1948. Bilatcial fractional resection of\nfiontal coitex for the treatment of psychoses. J. Nero Ment. Dis , 107\n411-429.\n\nHebb, D. O. 1937a. The innate organization of visual activitj’; I. Per-\nception of figures by rats reared in total darkness. J Genet. Psychol.,\n51, 101-126.\n\nHebb, D. O. 1937 /j. The innate organization of visual actii'ity: II.\nTransfer of lesponso in the discrimination of brightness and size by rats\nleared in total darkness. 1. Comp. Psychol., 24, 277-299.\n\nHclrb, D. O. 1938a. Studies of the organization of behavior: I. Behavior\nof the rat in a field orientation. J Comp. Psychol., 25, 333-352.\n\nHeblr, D. O. 1938/5. Studies of the organization of behavior: II.\nChanges in the field orientation of the rat after cortical destruction.\nJ. Comp. Psychol., 26, 427-444.\n\nPlebb, D. O. 1939. Intelligence in man after large removals of cerebral\ntissue; report of four left frontal lobe cases. ]. Gen. Pstfchol., 21, 73-87.\n\nHebb, D. O. 1942a. The effect of early and late brain injury upon test\nscores, and the nature of normal adult intelligence. Proc. Anier. Phil.\nSoc., 85, 275-292.\n\nPlebb, D. O. 1942 / 5 . Verbal test material independent of special vocabu-\nlary difficulty. J. Educ. Psychol., 33, 691-696.\n\nHebb, D. O. 1945a. The forms and conditions of chimpanzee anger.\nBidl. Canad. Psychol. Ass., 5, 32-35.\n\nHebb, D. O. 1945/5. Man’s frontal lobes: A critical review. Arch.\nNeurol. Psychiat., 54, 10-24.\n\nHebb, D. O. 1946a. Emotion in man and animal; an analysis of tire\nintuitive irrocesses of recognition. Psychol. Reo., S3, 88-106.\n\nPlebb, D. O. 1946/5. On the nature of fear. Psychol. Rev., S3, 259-276.\n\nHebb, D. O. 1947. Spontaneous neurosis in chimpanzees: tlieoretitJal re-\nlations with clinical and experimental phenomena. Psychosom. Med.,\n9, 3-16.\n\n\n\n310 Bibliography\n\nHebb, D. O., and Foord, E. N. 1945, Errors of visual recognition and\nthe nature of die ti-ace. J. Exp. Psychol, SS, 335-348.\n\nHebb, D. 0., and Morton, N. W. 1943. The McGill Adult Comprehen-\nsion Examination; \"Verbal Situation” and “Picture Anomaly” Series.\nJ. Educ. Psychol, 34, 16-25.\n\nHebb, D. 0., and Morton, N. W. 1944. Note on dio measurement of\nadult intelligence. J. Gen. Psychol, 30, 217-223,\n\nIlebb, D. O., and Penfleld, W. 1940. Human behai'ior after extensive\nbilateral removal from the frontal lobes. Arch. Neurol Psychial., 44,\n421-438.\n\nHebb, D. O., and Riesen, A. H. 1943. The genesis of irrational fears.\nBull Canad. Psychol Ass., 3, 49-50.\n\nHebb, D. O., and Williams, K. 1941 E.-sperimental control of cues de-\ntermining the rat’s orientation. Bull. Canad. Psychol. Ass., 1, 22-23.\n\nHebb, D. 0., and Williams, K. 1946. A mediod of rating animal intel-\nligence. J. Gen, Psychol, 34, 59-65.\n\nHerrick, C. J. 1929. The thinking machine. Chicago: Univ. Chic. Press.\n\nHilgard, E. R., and Marquis, D. G. 1940. Conditioning and learning.\nNew York; Appleton-Century.\n\nHoagland, H 1947. Enzyme kinetics and the dynamics of behavior.\nJ. Comp. Physiol. Psychol, 40, 107-127.\n\nHoagland, H., Malamud, W., Kaufman, I. C., and Pincus, G. 1946.\nChanges in the electroenceiihalogram and in the e.xcretion of 17-ketoster-\noids accompanying electi'o-shock dierapy of agitated depression. Psy-\nchosotn. Med., 8, 246-251.\n\nHobbs, G. E. 1941. Mental disorder in one of a pair of identical twins.\nAmcr. J. Psychial, 98, 447-450.\n\nHobhouse L. T. 1915. Mind in evolution. 2nd Ed. London: Mac-\nmillan.\n\nHovland, C. I. 1930. \"Inhibition of reinforcement” and phenomena of\nexperimental extinction. Pjoc. Nat. Acad. ScL, Wash., 22, 430-433.\n(Quoted by Hilgard and Maiquis, 1940, p. 146.)\n\nHovland, C. I. 1937. The generalization of conditioned responses. II.\nThe sensory generalization of conditioned responses with varying inten-\nsities of tone. J. Genet. Psychol, SI, 279-291.\n\nHuU, C. L. 1934. The concept of the habit-family hierarchy and maze\nlearning. Psychol Rev., 41, 33-54; 134-152.\n\nHull, C. L. 1943. Principles of behavior: an introduction to behavior\ntheory. New York: Appleton-Century.\n\nHuU, C. L. 1945. The discrimination of stimulus configurations and the\nhypothesis of afferent neural interaction. Psychol Rev., 52, 133-142.\n\nHumphrey, G. 1940. The problem of the direction of thought. Brit. J.\nPsychol, 30, 183-198.\n\nHunt, J. McV. 1941, The effects of infant feeding-frustration upon adult\nhoarding in tire albino rat. J. Abn. Soc. Psychol, 36, 338-360.\n\n\n\n311\n\n\nBibliography\n\nHunter, W. S. 1934. Learning: IV. Experimental studies of learning.\nIn Murchison, C., Handbook of general experimental psychology.\nWorcester, Mass.: Clark Univ. Press, pp. 497-570.\n\nJackson, T. A. 1942. Use of the stick as a tool by young chimpanzees.\nJ. Comp. Fsijchol., 34, 22,3-235.\n\nJacobsen, C. F., Jacob.sen, M. M., and Yoshioka, J. G. 1932. Derelop-\nmcnt of an infant chimpanzee during her first year. Comp. P.itichol.\nMonog., 9, 1-94.\n\nJames, W. 1910. Principles of psychology. New York: Holt.\n\nJasper, H. H. 1937. Electrical signs of coibcal activity'. Psychol. Bull.,\n34, 411-481.\n\nJasper, H. H. 1941. Electroencephalography. In Penfield, W., and\nErickson, T. C., Epilepsy and cerebral localization. Springfield: Thomas,\npp. 380-454.\n\nJa.sper, H. H., and Fortuyn, J. D. 1946. E.xperimental studies on the\nfunctional anatomy of petit mal epilepsy. Pnhl. Ass. Res. Nerv. Meat.\nDi.s., 26, 272-298.\n\nJeiferson, G. 1937. Removal of right or left frontal lobes in man. Brit.\nMed. I., 2, 199-206.\n\nJersild, A. T., and Holmes, F. B. 1935. Children’s fears. New Yoik:\nTeach. Coll. Bui. Publ.\n\nJolhffc, N. 1942. The neurop-sychiatric manifestations of vitamin defi-\nciencies. I. Mi. Sinai Hasp., 8, 658—667.\n\nJones, H. E., and Gonrad, H. S. 1933. The growth and decline of intel-\nligence: a study of a homogeneous group between the ages of ten and\nsixty. Genet. Psychol. Monog., 13, No. 3.\n\nJones, H. E., and Jones, M. C. 1928. A study of fear. Childhood Educ.,\n5, 136-143.\n\nJones, M. G. 1933. Emotional development. In Murchison, C., A hand-\nbook of child psychology. 2nd Ed. Worcester, Mass.: Clark Univ. Press,\npp. 271-302.\n\nKappers, C. U. A., Huber, G. C., and Crosby, E. C. 1936. The compara-\ntive anatomy of the nervous system of vertebrates, including man. New\nYork: Macmillan, Vol. I.\n\nKarnosh, L. J., and Gardner, W. J. 1941. An evaluation of the physical\nand mental capabilities following removal of the right ceiebral henii-\nsiihere. Cleveland Clin. Quart., 8, 94-106.\n\nKcnnard, M. A. 1939. Alterations in response to visual stimuli following\nlesions of frontal lobe in monkeys. Arch. Neuwl. Psychiat., 41, 1153-\n1165.\n\nKennard, M. A., and Ectors, L. 1938. Forced circling in monkeys follow-\ning lesions of the frontal lobes. J. Neurophysiol., 1, 45-54.\n\nKennedy, F., and Wolf, A. 1936, The relationship of intellect to speech\ndefect in aphasic patients. I. New. Ment. Dis., 84, 125-145; 293-3] 1.\n\nKeschner, M,, Bender, M., and Strauss, I. 1938. Mental symptoms asso-\n\n\n\n312 Bibliography\n\nciated with brain tumoi; a study of 530 verified cases. J. Arner. Med.\nAss., no, 714-718.\n\n'Kinder, E. F. 1927. A study of die nest-building actiiity of tlie albino\nrat. J. Exp. Zool, 47, 117-161.\n\nKinsey, A. C., Pomeroy, W. B., and Martin, C. E. 1948. Sexual behavior\nin the human male. Philadelphia: Saunders.\n\nKlebanoff, S. G. 1945. Psychological changes in organic brain lesions and\nablations. Psychol. Bull., 42, 585-623.\n\nKleitman, N. 1939. Sleep and wakefulness. Chicago: Univ Chic. Piess.\n\nKlineberg, O. 1940. Social psychology. New York’ Holt.\n\nKoffka, K. 1924. The growth of the mind. New Yoik: Haicourt, Brace.\n\nKoffka, K. 1935. Principles of Gestalt psychology. New York: Haroourt,\nBrace.\n\nKohler, \\V. 1925. The mentality of apes. New York: Harcourt, Brace.\n\nKohler, W. 1929. Gestalt psychology. New York: Liveiight.\n\nKohler, W. 1940. Dynamics in psychology. New York: Liveright.\n\nKohler, W., and Wallach, H. 1944. Figural after-effects: an investigation\nof visual proce.sses. Proc. Amer. Phil Soc., 88, 269-357.\n\nKrechei'.sky, I 1932. “Hypotheses” versus “chance” in the pre-solution\nperiod in sensory discrimination-learning. Univ. Calif. Publ. Psychol,\ne, 27-44.\n\nKrechevsky, I. 1938. An experimental investigation of the principle of\nproximity in the vi.sual perception of the rat. J. Exp. Psychol, 22,\n497-523.\n\nKiibitschek, P. E. 1928. The symptomatology of tumoi s of the frontal\nlobe based on a series of twenty-two cases. Arch. Neurol. Psychiat., 20,\n559-579.\n\nLandis, C. 1947. A modern dynamic psj'chology. J. Comp. Physiol\nPsychol, 40, 135-141.\n\nLandis, C , and Hunt, W. A. 1932. Adrenalin and emotion. Psychol.\nRev., 89, 467-485.\n\nLashley, K. S. 1929<i. Brain mechanisms and inlelUgence: a quantitative\nstudy of injuries to the brain. Chicago: Univ. Chic. Press.\n\nLashley, K. S. 19295. Nervous mechanisms in learning. In Murchi-\nson, C., The foundations of experimental psychology. Worcester: Clark\nUniv. Press, pp. 524r-563.\n\nLashley, K. S. 1930. Basic neural mechanisms in behavior. Psychol.\nRev., 37, 1-24.\n\nLashley, K. S. 1934. The mechanism of vision. VIII. The projection of\nthe retina upon the cerebral cortex of the rat. J. Comp. Neurol, 60,\n57-79.\n\nLashley, K. S. 1937. Functional determinants of cerebral localization.\nArch. Neurol Psychiat., 38, 371-387.\n\nLashley, K. S. 1938a. Experimental analysis of instinctive behavior.\nPsychol Rev,, 43, 445-471.\n\n\n\n813\n\n\nBibliography\n\nLashley, K. S. 1938ir, The mechanism of vision: XV. Preliminaiy studies\nof the lat’s capacity for detail vision. ]. Gen. Psychol., 18, 12.3-193.\n\nLashley, K. S. 1938c. The thalamus and emotion. Psychol. Rev , 45,\n42-61.\n\nLasiiley, K. S. 1941. Patterns of cerebral integration indicated by the\nscotomas of migraine. Arcli. Neurol. Psychiat., 46, 331-339.\n\nLashloy, K. S. 1942a. The problem of cerebral organization in \\ision.\nIn Kluver, H., Visual mechanisms. Biol. Stjmpos., 7, 301-322.\n\nLashley, K. S. 1942h. An examination of the “continuity dieory” as ap-\nplied to discnminatimi learniag. 1. Gen. Psychol, 26, 241-26.5.\n\nLashley, K. S. 1944. Studies of cerebral’ function in learning; XIII.\nApparent absence of trahscordcal association in maze learning. J. Comp.\nNeurol, SO, 257-281.\n\nLashley, K. S., and Clark, G. 1946. The eytoarchitecture of the cerebral\ncoitex of Ateles: a critical examination of architectonic studies. J. Comp.\nNeurol, 85, 223-306.\n\nLashley, K. S., and Wade, M. 1946. The Pavlovian theory of generaliza-\ntion. Psychol. Rev., 53, 72-87.\n\nLeeiier, R. W. 1935. A study of a neglected portion of tlie field of\nlearning— die development of sensory organization. J. Genet. Psychol, 46,\n41-75.\n\nLeeper, R. W. 1948. A motivational theory of emotion to replace “Emo-\ntion as di.soiganized response.” Psychol. Rev., 55, 5-21.\n\nLehmann, J. E. 1937a. The effect of changes in the potassium-calcium\nbalance on die action of mammalian A nerve fibers. Amer. J. Physiol.\n118, 613-619.\n\nLehmann, J. E. 1937h. The effect of changes in pH on the action of\nmammalian A nerve fibers. Amer. J. Physiol, 118, 600-612.\n\nLevine, J. 1945a. Studies in the interrelations of central nervous struc-\ntures in binocular vision: I. The lack of bilateral transfer of I'isual dis-\ncriminadve habits acquired monocularly by the pigeon. J. Geiret. P.sy-\nchol, 67, 105-129.\n\nLei’ine, J. 1945h. Studies in the interrelations of central nervous sti-uc-\ntures in binocular vision; II. The condidons under which intcrocular\ntransfer of discriminative habits takes place in die laigeon. J. Genet.\nPsychol., 67, 131-142.\n\nLewin, K. 1938. Will and needs. In Ellis, W. D., A source book of\nGestalt psychology. London: Kegan Paul, Trench, Tiubner, pp. 283-299.\n\nTibet, B., and Gerard, R. W. 1039. Control of the potential rhythm of\nthe isolated frog brain. J. Neurophysiol, 2, 153-169.\n\nLiddell, H. S. 1938. The experimental neurosis and die pioblem of\nmental disorder. Amer. J. Psychiat, 94, 1035—1041.\n\nLiddell, H. S. 1944. Animal behavior studies bearing on die problem of\npain. Psychosom. Med., 6, 261-263.\n\nLorente de No, R. 1938a. Synaptic stimulation of motoneurons as a local\nprocess. 1. Neurophysiol, 1, 195-206.\n\n\n\n314 Bibliography\n\nLorente de N6, R. 1938&. Analysis of ihe activity of the chains of inter-\nnuncial neurons. J. Neurophyi,iol , 1, 207-244.\n\nLorente de No, R. 1939. Tiansmission of impulses through cranial motor\nnuclei. I. NeuropliysioL, 2, 402-464.\n\nLorente de No, R. 1943. Cerebral cortex; architecture. In Fulton, J. F.,\nPhysiology of the nervous system. 2nd Ed. New York. Oxford Univ.\nPress, pp. 274-301.\n\nLorenz, K. 1935, Dei Kimipan in der Umwelt des Vogels. J. Ornith., S3,\n137-213; 289-413.\n\nLoucks, R. B. 1935. Experimental delimitfition of neural structui'es essen-\ntial for learning; The attempt to condition striped muscle responses with\nfaradization of the sigmoid gyii. J. Psychol., 1, 5-44.\n\nLoucks, R. B. 1938. Studies of neural Structuies essential for learning.\nII. The conditioning of salivary and striped muscle responses to faradi-\nzation of cortical sensory elements, and tire action of sleep upon such\nmechanisms. J. Comp. Psychol., 25, 315-332.\n\nLuckhardt, A. B., and Carlson, A. J. 1915. Contributions to the physi-\nology of the stomach. XVII. On flie chemical control of llie gasti'ic\nhunger mechanism. Amer. J. Physiol, 36, 37-46.\n\nMcBiide, A. F., and Hebb, D. O. 1948. Behavior of the captive bottle-\nnose dolphin, Tursiops truncatus. J. Comp, Physiol. Psychol, 41, 111-\n123.\n\nMcCulloch, T. L., and Haslerud, G. M. 1939. Affective responses of an\ninfant chimpanzee reared in isolation from its land. J. Comp. Psychol,\n28, 437-445.\n\nMcCullocIi, W. S. 1944a. Cortico-cortical connections. In Buoy, P., The\npiecential motor cortex. Uibana, 111.: Univ. Illinois Press, pp, 213-242.\n\nMcCullo«h, W. S. 1944&. The functional oiganization of the cerebral\ncortex. Physiol Rev., 24, 390—407.\n\nMcGeoch, J, A 1942. The psychology of human learning. New York:\nLongmans, Green.\n\nMaier, N. R. F., and Schneiila, T. C. 1935. Principles of animal psy-\nchology. New York; McGraw-Hill.\n\nMaishall, W. H., and Talbot, S. A. 1942. Recent evidence for neuial\nmechanisms in vision leading to a general theory of sensory acuity. In\nKliiver, H., Visual mechanisms. Biol. Symp., 7, 117-164.\n\nMasserman, J. H. 1942. The hypotlialamus in psychiati'y. Amer. 1.\nPsychiat., 98, 633-637.\n\nMasserman, J. H. 1943. Behavior and neurosis' An experimental psycho-\nanalytic approach to psychobiologic principles. Chicago: Univ, Chic.\nPress.\n\nMatthews, R. S. 1938. Pellagra and nicotinic acid. 1. Amer. Med. Ass.,\nITl, 1148-1153.\n\nMiller, G. A. 1947. The masking of .speech. Psychol Bull, 44, 105-129.\n\nMiner, J. B. 1905. A case of vision acquired in adult life. Psychol. Rev.\nMonog. Stippl, 6, No. 5, 103-118.\n\n\n\n315\n\n\nBibliography\n\nMixtei, W. J., Tillotson, K. J., and Wies, D. 1941. Reports of partial\nfrontal lobectomy and frontal lobotoniy performed on three patients; one\nchronic epileptic and two cases of chronic agitated depression. Psy-\nchosom. Med., 3, 26-37.\n\nMoigan, C. T. 1943. Physiological psychology. New York; McGiaw-Hill.\n\nMonson, R. S., and Dempsey, E. W. 1943. Mechanism of thalamocortical\naugmentation and repetition. Amer. J. Physiol, 138, 297-308.\n\nMoss, F. A. (Ed.). 1942. Compaiatioe psychology. Rev. Ed. New York;\nPrentice-Hall.\n\nMowrer, O. H. 1941. MotivaMon and learning in relation to the national\nemergency. Psychol. Bull, 38, 421-431.\n\nMowrer, O. H., and Mowrer, W. M. 1938. Enuresis— a mctliod for its\nstudy and treatment. Amer. •?. Orlhopsychiat., 8, 436-A5Q.\n\nMuenzinger, K. F. 1934. Motivation in learning. 1. Electric shock for\ncorrect response in tlie visual discrimination habit. J. Comp, Psychol,\n17, 267-277.\n\nMurphy, J. P., and Gellhorn, E. 1945. Further investigations on dien-\nceiihalic-cortical relations and their significance for the problem of emo-\ntion. J, Neurophysiol, 8, 431-447.\n\nNafe, J. P. 1934. The pressure, pain, and temperature senses. In Murchi-\nson, C., Handbook of general experimental p.sychology. Worcester,\nMass.: Clark Univ. Press, pp. 1037-1087.\n\nNauta, W. ,T. H. 1946. Hypothalamic regulation of sleep in rats; An ex-\nperimental study. J. Neurophysiol, 9, 285-316.\n\nNeet, C, C. 1933. Visual pattern discrimination in the Macacus rhesus\nmonkey. 1. Genet Psychol, 43, 163-196.\n\nNeff, Walter S. 1938. Socioeconomic status and intelligence: a ciitical\nsurvey. Psychol. Bull, 35, 727—757.\n\nNichols, I., and Hunt, J. McV. 1940. A case of partial bilateral frontal\nlobectomy: a psychopathological study. Amer, J. Psy chiat, 96, 1063-\n1083.\n\nNissen, H. W., Machover, S., and Kinder, E. F. 1935. A study of per-\nformance tests given to a group of native African negro children. Brit.\nJ. Psychol, 23, 308-355.\n\nPavlov, I. P. 1927. Conditioned reflexes. Oxford: Humphrey Milford.\n\nPavlov, I. P. 1928. Lectures on conditioned reflexes. New York; Inter-\nnational,\n\nPavlov, I. P. 1932. The reply of a physiologist to psychologists. Psychol.\nBeo., 39, 91-120.\n\nPenfield, W., and Jasper, H. 1946. Highest level seizures. Res. Publ.\nAss. Nero. Ment. Dis., 26, 252—271.\n\nPennington, L. A, 1938. The function of the brain in auditory loc^iza-\ntion. IV. Metliod of training and control experiments. J. Comp\nPsychol, 25, 195-211.\n\nPilgrim, F. J,, and Patton, R. A. 1947. Patterns of self-selection of puri-\n\n\n\n316 Bibliographij\n\nBed dietary components by the lat. 7. Comp. Physiol. Psychol., 40,\n343-348.\n\nPillsbury, W. B. 1913. “Fluctuations of attention” and the refractory\nperiod. 7. Phil. Psychol. Sci. Meth., 10, 181-185.\n\nPolyak, S. L. 1941. The tetina. Chicago: Univ. Chic. Press.\n\nPostman, L. 1947. The histoiy and present status of the law of effect.\nPsychol. Bull., 44, 489-563.\n\nPratt, C. C. 193^). The logic of modem psychology. New York: Mac-\nmillan.\n\nPrentice, W. C. H. 1946. Operationism aod psychological tlreoiy: a note.\nPsychol. Reo., S3, 247-249.\n\nProsser, C. L. 1934. Action potentials in the nervous system of the cray-\nfish: I. Spontaneous impulses. 7. Cell. Comp. Physiol., 4, 185-209.\n\nRichter, C. P., Holt, L. E., and Barelare, B. 1938. Nutritional require-\nments for normal growth and reproduction in rats studied by tlie self-\nselection method. Amer. 7. Physiol., 122, 734-744.\n\nRiddoeh, G. 1941. Phantom limbs and body shape. Brain, 64, 197-222.\n\nRiesen, A. H. 1947. The development of visual perception in man and\nchimpanzee. Science, 106, 107—108.\n\nRiesen, A. H., and Nissen, H. W. 1942. Non-spatial delayed response by\nthe matching technique. 7. Comp. Psychol, 34, 307-313.\n\nRoethlisburger, F. J., and Dickson, W. J. 1939. Management and the\nworker. Cambridge: Han'ard Univ. Press.\n\nRomano, J., and Coon, G P. 1942. Physiologic and psychologic studies\nin spontaneous hypoglycemia. Psychosom. Med., 4, 283-300.\n\nRowe, S. N. 1937. Mental changes following the removal of the right\ncerebral hemisphere for brain tumor. Amer. 1. Psychiat., 94, 605-614.\n\nRubin, 1921. Visuell wahrgenommene Figuren: Studien in psycho-\nlogischer Analyse. Teil I. Berlin: Gyldendalske Boghandel.\n\nRylander, G. ,1939. Personality changes after operations on the frontal\nlobes: a clinical study of 32 cases. London: Humphrey Milford.\n\nSaphir, W. 1945. Chronic hypochlorcmia simulating psyohoneurosis.\n7. Amer. Med. Ass., 129, 510-512.\n\nSchneirla, T. C. 1948. Psychology, comparative. Encycl. Brit.\n\nScott, W. W., Scott, C. C., and Luckhardt, A. B. 1938. Observations on\nthe blood sugar level before, during, and after hunger periods in humans.\nAmer. I. Physiol, 123, 243-247.\n\nSenden, M. v. 1932. Baum- und Gestaltauffassung bei operierten Blind-\ngeborenen vor und nach der Operation.. Leipzig: Barth.\n\nSherrington, C. S. 1906. Integrative action of the nervous system. New\nYork; Scribner.\n\nSherrington, C. S. 1925. Remarks on some aspects of reflex inhibition.\nP!“c. Roy. Soc., 97B, 519-545.\n\nSherrington, C. S. 1941. Man on his nature. New York: Macmillan.\n\nSkinner, B. F. 1938. The behavior of organisms: an experimental analysis.\nNew York; Appleton-Centuiy.\n\n\n\nBibliography 317\n\nSmith, D. E. 1939. Cerebral localization in somestlielic discrimination in\nthe rat. J. Comp. Psychol., 28, 161-188.\n\nSmitli, K. U. 1936. Visual discrimination in the cat; III. The lelative\nefleot of paired and unpaired stimuli in tlie discriminative behavior of\nthe cal. J. Genet. Psychol., 48, 29-57.\n\nSmith, W. K. 1945. The functional significance of tlie rostral cingular\ngyius as revealed by its responses to electrical excitation. J. Neuro-\nphysiol., 8, 241-255.\n\nSpearman, C. 1927. The abilities of man. New York: Macmillan.\n\nSpence, K. W. 1938. Gradual versus sudden solution of discrimination\nproblems by chimpanzees. J. Comp. Psychol., 25, 213—224.\n\nSpence, K. W. 1940. Continuous versus non-continuous interpretations of\ndiscrimination learning. Psychol. Rev., 47, 271-288.\n\nSperiy, R W. 1943. Visuomotor cooidination in the newt {Triturus\nviridcscens) after regeneration of the optic neive. J. Comp. Neurol., 79,\n33-55.\n\nSpeiry, R. W. 1947. Effect of crossing nerves to antagonistic limb mus-\ncles in the monkey Arch. Neurol. Psychiat., 58, 452—473.\n\nSpiegel, E. A , Millei, H. R., and Oppenheimer, M, J. 1940. Forebrain\nand rage reactions. /. Nouiophysiol., 3, 539-548.\n\nSpie.s, T. D,, Aring, C. D., Gelperiii, J., and Bean, VV. B. 1938. The\nmental symptoms of pellagra: Their relief with nicotinic acid. Amer. J.\nMed. Sci., 190, 461-475.\n\nSpragg, S. D S. 1940. Morphine addiction in chimpanzees. Comp.\nPsychol. Monog., 15, No. 7.\n\nStoddard, G. D., and Wellman, B. L. 1940. Environment and the IQ.\nYearb. Nat. Soc. Stud. Edttc., 39 (I), 405—442.\n\nStookey, B., Scarff, J., and Teitelbaum, M. 1941. Frontal lobettomy in\nthe treatment of brain tumors. Ann. Surg., 113, 161-169.\n\nSwank, R. L., and Marc'hand, W. E. 1946. Combat neuroses: develop-\nment of combat exhaustion. Arch. Neurol. Psychiat., 55, 236-247.\n\nThorndike, E L. 1931. Human learning. New York; Century.\n\nThurstone,''L. L. 1935. The vectors of mind. Chicago; Univ. Chic. Press.\n\nTinbergen, N. 1942. An objectivistic study of die innate behavior of\nanimals. Bibl. Biotheoret., Leiden, 1, 39-98.\n\nTinklepaugh, O. L. 1928. An experimental study of representative fac-\ntois in monkeys. J. Comp. Psychol., 8, 197-236.\n\nTitchener, E. B. 1920. Notes from the psychological laboratory of Cornell\nUniversity. Amer. J. Psychol , 31, 212—214.\n\nTolman, E. C. 1932, Purposive behavior in animals and men. New\nYork; Century.\n\nTryon, R. C. 1939. Studies in individual differences in maze learning: VI.\nDisproof of sensory components; experimental effects of stimulus varia-\ntion. J. Comp. Psychol, 28, 361-415.\n\nValentine, C. W. 1930. The innate bases of fear. J. Genet. Psychol, 37,\n394-419.\n\n\n\n318 Bibliography\n\nWalker, A. E., and Weaver, T. A. 1940. Ocular movements from the\noccipital lobe in tlie monkey. J. Neurophysiol., 3, 353-357.\n\nWatson, J. B. 1924. Behaviorism. New York; Norton.\n\nWatts, J. W., and Freeman, W. 1946. Psychosurgery for the relief of in-\ntractable pain. J. Int. Coll. Surg., 9, 679-683.\n\nWechsler, D. 1939. The measurement of adult intelligence. Baltimore;\nWilliams and Wilkins.\n\nWeddel, G., Sinclair, D. G., and Feindel, W. H. 1948 An anatomical\nbasis for alterations in quality of pain sensibility. J. Neurophysiol., 11,\n99-109.\n\nWeisenhuig, T., and McBride, K. E. 1935. Aphasia: a clinical and psy-\nchological study. New York: Commonwealth Fund.\n\nWeisenburg, T., Roe, A., and McBride, K E. 1936. Adult intelligence:\na psychological study of test performances. New York: Commonwealth\nFund.\n\nWeiss, P. 1941a. Autonomous versus reflexogenous activity of the cenh'al\nnervous system. Proc. Amer. Phil. Soc., 84, 53-64.\n\nWeiss, P. 19416. Nerve patterns: The mechanics of nerve growth.\nGrowth {Third Growth Symposium), S, 163-203.\n\nWerner, H., and Strauss, A. 1939. Types of visuo-motor activity in their\nrelation to low and high performance ages. Proc. Amer. Ass. Ment.\nDefic., 44, 163-168.\n\nWilson, G., and Rupp, C. 1947. Present trends in the practice of neu-\nrology. J. Amer. Med. Ass., 133, 509-511.\n\nWolf, E , and Zerrahn-Wolf, G, 1937. Flicker and lire reactions of bees\nto flowers, J, Gen. Physiol., 20, 511-518.\n\nWolf, G, A., and WoW, H. G. 1946. Studies on the nature of certain\nsymptoms associated with cardiovascular disorders. Psychosom. Med., 8,\n293-319.\n\nWolff, H. G. '1943. Emotions and gastric function. Science, 98, 481^84.\n\nWolff, H, G., and Hardy, J. D. 1947. On the nature of pain. Physiol.\nRev., 27, 167-199.\n\nWoodrow, H. 1927. The effect of type of training on transference.\nJ. Educ. Psychol., 18, 160-171.\n\nWoodworth, R. S, 1921. Psychology. New York: Holt.\n\nWoodworth, R.'S. 1938. Experimental psychology. New York: Holt.\n\nWords, H., Stein, M. H., and JoUiffe, N. 1942. Fiber dissociation in\nperipheral neuropathy. Arch. Int. Med., 69, 222-237.\n\nYerkes, R. M. 1916. The mental life of monkeys and apes: a study of\nideational behavior. Behavior Monog., 3, No. 1.\n\nYoung, P. T. 1941. The experimental analysis of appetite. Psychol. Bull.,\n3&, 129-164.\n\nYoung, P. T. 1944. Studies of food preference, appetite and dietary habit.\nI. Running activity and dietary habit of the rat in relation to food pref-\nerence. 1. Comp. Psychol., 37, 327-370.\n\n\n\nBibliography S19\n\nZangwill, O. L. 1937. A study of tlie significance of attitude in lecogni-\nlion. Brit, J. 28, 12-17.\n\nZener, K. 1937 The significance of heliavior accompanying conditioned\nsaliMiry secretion foi theories of tlie conditioned response. Aincr. J.\nPsychol., 50, 384—403.\n\nZollingci, R. 198.0. Renioral of left ccreliial liemispheie. lepuit of a case\nArch, Neurol, Psijchiat., 34, 1055-1064."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "All 15 bibliography pages are indexed, from Adams through Zollinger. The archival OCR is retained as reference text; a new Windows OCR pass is also preserved for every page with image and text hashes. OCR spelling and punctuation errors remain; citation matches are reviewed separately."
      ],
      "BibliographyEdition": "1949 first edition",
      "OcrPages": [
        {
          "PdfPage": 328,
          "Text": "Bibliography\nAdams, D. K. 1929. Experimental studies of adaptive behavior in cats.\nComp. Psychol. Monog.. 6, 'No. 1.\nAdrian, E. D. 1981. Potential changes in the isolated nep.'0us system of\nDytiscus nunginaZis. J. Physiol., 72, 182—151.\nAdrian, E. D. 1984. Electncal activity of the nervous system. Arch,\nNeurol. Psychiat., 82, 1125—1186.\nAdrian, E. D , and BuytendiJk, F. J. J. 1981. Potential changes in the iso-\nlated brain stem of the goldfish. J Physiol, 71, 121—135.\nAdrian, E. D., and Matthews, B. H. C. 19.34. The interpretation of po-\ntential waves in the cortex. J. Physiol„ 81, 440—471.\nAllen, C. , and Broster, L. R. 1945. A further case of paranoid psychosis\nsuccessfully treated by adrenalectomy. Brit. Med. J. , No. 4402, 696—698.\nAllport, G W. 1946. Effect: a secondary principle of learning. Psychol.\nRec., 53, 835-347.\nAnderson, J. E. 1939. The limitations of infant and preschool tests in the\nmeasurement of intelligence. J. Psychot., 8, 851—879.\nArvanitnki, A. 1942. Effects evoked in an axon by the activity of a con-\ntiguous one. J. Neurophysiol., 5, 89—108.\nBard, P. 1984. On emotional expression after decortication with some re-\nmatks on certain theoretical views. Psychol. Ren. , 41, 309—339.\nBard, P. 1942. Neural mechanisms in emotional and sexual behavior,\nPsychosom. Med, 4, 171—172.\nBartley, S. H. , and Bishop, G. H. 1988. Factors determining the form of\nthe electrical response from the optic cortex of the rabbit. Amer. J.\nPhysloZ., log, 173-184.\nBartley, S. I-1., and Chute, E. 1947. Fatigue and impairment in man.\nNew York: McGraw-Hill.\nBeach, F. A. 1987. The neural basis of innate behavior. I. Effects of\ncortical lesions upon the maternal behavior pattern in the rat. J. Comp.\nPsychol., 24, 898-439.\nBeach, F. A. 1989. The neural basis of innate behavior. III. Compari-\nson of learning ability and instinctive behavior in the rat.\nJ. Comp.\nPsychot., 28, 225-262.\nBeach, F. A. 1942. Analysis of factors involved in the arousal, mainte-\nnance and manifestation of sexual excitement in male animals. Psy-\nchosom. Med, 4, 173-198.\nBench, F. A. 1947a. A review of physiological and psychological studies\nof sexual behavior in mammals. Physiol. Reu, 27, 240—807.\n805",
          "Receipt": {
            "TextSha256": "3A3BE38348465EB34FE7D7EACE751504CFD4E3621053D4CFDC535AF9A60B176D",
            "ProcessedAtUtc": "2026-09-16T23:48:46.1737707Z",
            "ImageSha256": "7B328532D8E0EF4B92CC7A07DC552A14651B9F0760D8129DF5765E34F52CF6C2",
            "Language": "en-US",
            "Lines": 41,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000328.ocr.txt",
            "Image": "hebb-refs-000328.png"
          }
        },
        {
          "PdfPage": 329,
          "Text": "806\nBibliography\nBeach, F. A. 1947b. Evolutioua1Y' changes m the physiological control Of\nmating behavior in mammals. Psychol. Ret., 54, 297—815.\nBeach, F. A. 1948. Hormones and behauor. New York: Hoeber.\nBellak, L. , and Willson, E. 1947. On the etiology of dementia praecox\nJ. Ne,u. Ment. Dis., 105, 1-24.\nBellows, R. T. 1989. Time factors water drinking m dogs. Amer. J.\nPhysiol., 125, 87-97.\nBirch, H. G. 19*. The relation of previous experience to inswhtful\nproblem-solving. J. Comp. Psychol., 88, 867—883.\nBishop, C. H. 1946. Nerve and synaptic ccy-lducuon. Ann. Rec. Physiol.,\n8, 355-374.\nv. Bonin, C, Carol, H. Wv, and McCulloch, W. S. 1942. The functional\norganization of the occipital lobe, In kluver, H., Visual mechanisms.\nBiol. Symp., 7, 165-192.\nBoring, E. G. 1916. Cutaneous sensation after nave-dii']slon. Quart. J.\nExp. Physiot., 10, 1-95.\nBoring, E G. 1980. A new ambiguous figure. Amer. J. Psychot., 42,\n444—445.\nBoring, E. C. 1938, The physical dimensions of consciousness. New\nYork: Century.\nBoring, E. G. 1946. Mind and mechanism. Amer. J. Psychol, 59, 173—\n192.\nBousfield, W. A. 1985. Quantitative indices of lhe effects of fasting on\neating-behavior. J. Genet. Psychol., 46, 476—479.\nBowman, K. M.\n1935. Psychoses With permcious anemia, Amer. J.\nPsychiat., 92, 371-396.\nBowman, K. M. 1946. Modern concept of the neuroses. J. Amer. Med.\nAssoc., -182, 555-557.\nBridgman, C. S., and Smith, K. U. 1945. Bilateral neural integmtion in\nvisual perceptE)n after section of the corpus callosum. J. Comp. Neuro!.,\n88, 57-68.\nBronk, D. W. 1939. Synaptic mechanisms in sympathetic ganglia.\nJ. Neurophysiol., 2, 880—401.\nBrown, Warner. 1982. Spatial integrations in a human maze. Unit'.\ncalif. Publ. Psychol., 5, 123-134.\nBruetsch, W. L. r 1947. Rheumatic brain disease: late sequel of rheumatic\nfever. J. Amer. Med. Assoc., 134, 450—454.\nCarlson, A. J. 1916. The control of hunger in health and disease. Chi-\ncago: Univ. Chic. Press.\nCarmichael, L., Hogan, H. P., and Walter, A. A. 1932. An experimental\nstudy of the effect of language on the reproduction of visually perceived\nform. J. Exp. Psychol., 15, 73—86.\nChaffGler, A. R. 1984. Beauty and human nature. New York: Appleton-\nCentury.\nClark, G., and Lashley, K. S. 1947. Visual disturbances following frontal\nablations in the monkey. Anat. Rec., 97, 826.",
          "Receipt": {
            "TextSha256": "DC6AED31FE719C75C0512882A3E47D2671533A829B940FF3FCC7CCF6D17F32FD",
            "ProcessedAtUtc": "2026-09-16T23:48:46.3812008Z",
            "ImageSha256": "67DD61C34620FD71D8E23AEDBA57C328F475F05BD23723044185A420B12E437C",
            "Language": "en-US",
            "Lines": 49,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000329.ocr.txt",
            "Image": "hebb-refs-000329.png"
          }
        },
        {
          "PdfPage": 330,
          "Text": "Bibliography\n807\nCobb, S. 1944. Personality as affected by lesions of the brain. In Hunt,\nJ. McV., Personality and thc behavior dis07ders. New York: Ronald,\nVol. I, pp. 550—581.\nCobb, S., Cohen, M. E., and Badal, D. W. 1946. Capillaries of the nail\nfold in patients with neurocirculatory asthenia ( effort syndrome, an-sacty\nneurosis). Arch. Neurol. Psychiat , 56, 648—650.\nCohn, R. 1945. Electroencephalographic study of prefrontal lobotomy,\nArch. Neurol. Psychiat., 53, 283—288.\nCowles, J. T., and Nissen, H. W. 1937. Rewatd-expectancy in delayed\nresponses of chimpanzees, J. C oynp. Psychol., 24, 345—358.\nDandy, E. 1938. Physiologic studies following extirpation of the right\ncerebral hemisphere in man. Bull. Johns Hopkins Hosp., 53, 81—51.\nDaniel, S. , and Smith, K. 13. 1947. The sea-approach behavior of the\nneonate loggerhead turtle. J. Coynp. Physiol. Psvchol., 40, 418—420.\nDashiell, J. F. 1928. Are there any native emolions? Psychol. Rev., 35,\n819-827.\nDavison, C , and Demuth, E. L. 1945. Disturbances in sleep mechanism:\na clinico-pathologie study. II. Lesions at the corticodiencephahc level.\nArch. Neurol. Psychiat., 54, 241—255.\nDenker, P G. 1946. Results of treatment of psychoneuroses by the gen-\neral Inaclltioner: a follow-up study of 500 cases. N. Y. State I. Med.,\n'16, 2164-2166.\nDennis, W. 1984. Congenital cataract and unlearned behavior. J. Genet.\nPsychol., 44, 840-350.\nDennis, W. 1940. Infant reaction to restraint: an evaluation of Watson's\nlhecny. Trans. N. Y. Acad. sci., ser. 2., 2, No. 8, 202-218\nDenny-Brown, D. 1982. Theoretical deductions from the physiology of\nthe ccnebral cot tex. J. Neurol. psychopathol., 18, 52—67.\nDewan, J. G., and Owen, T. 1945. N'lcntal illness and the principles of\nmedicine. Caned. Med. Ass. J., 52, 349—357.\nDoll, E. A. 1983. Psychological significance of cerebral birth lesions.\nAmer. J. Psychol., 45, 444—452.\nDoll, E. A., Phelps, W. M., and Melcher, R. T. 1982. Mental deficiency\ndue to birth injuries. New York: Macmillan.\nDrew, G. C. 1988. The function of punishment in learning. J. Cenet.\nPsychol., 52, 257--267.\nDubner, H, H., and Gerard, R. W, 1939. Factors controlling brain po-\ntentials in the cat. J. Neurophysiol., 2, 142—152.\nDunlap, K. 1932. Habits: their making and remaking. New York:\nLiveright.\nEditors, Nutlition Reviews. 1944. Self-selection of diets. Nutrition Reu,\n2, 199-203.\nEgafia, E. , Johnson, R. E., Bloomfield, R., Brouha, L. , et al. 1942. The\neffects of a diet deficient in the vitamin B complex on sedentary men.\nAmor. J. PhusioZ., 187, 731-741.",
          "Receipt": {
            "TextSha256": "8AB864AB1D252D3CD7FA547BF2BDEF47FF77A4F18C71CDA146EF9262787D5A04",
            "ProcessedAtUtc": "2026-09-16T23:48:46.5869599Z",
            "ImageSha256": "E46EE05BDEE71CF1BF2502F888F2614AA443F9191A5610C2AA28D29E472A9846",
            "Language": "en-US",
            "Lines": 47,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000330.ocr.txt",
            "Image": "hebb-refs-000330.png"
          }
        },
        {
          "PdfPage": 331,
          "Text": "808\nBibliography\nErlanger, J. 1939. The iniüation of impulses in axons. J. Ncurophysiol.,\n2, 370-879.\nFerraro, A., Arieti, S., and Enghsh, W. H. 1945. Cerebral changes in\nthe eouxse of permmous anemia and their relationship to psychic symp-\ntoms. J. Neuropathol. Exp. Neu,ol., 4, 217-239.\nFields, P. E. 1982. Studies in concept formation. I. Comp. psychot.\nMonog., 9, No. 2.\nForbes, A. 1939. Problems of synaptic function. J. Neurophysiot., 2,\n465-472.\nFreeman, G. L. 1934. Introduction to pirustologicai psychology. New\nYork. Ronald.\nFreeman, W., and Watts, J. W. 1942. Psychosurgery: intelligence, emo-\ntion and social behavior following prefrcntal lobotomy for mental dis-\norders. Springfield: Thomas.\nFreeman, W. , and Watts, J. W. 1946. Psychosurgery. In Spiegel, E. A. ,\nProgress in neurology and psychiatry: an annual review. New York:\nGrune and Stratton, pp. 649—661.\nFuchs, W. 1920. Untersuchungen uber das Sehen der Hemianopker und\nIn Gelb, A., and Goldstein, K. , Psychologischen\nHemiamblyopiker: II.\nAnalysen hirnpathologischer Fauc. Leipzig: Darth, pp. 419—561.\nFulton, J. F. 1948. Physiology of the nertous system. 2nd Ed. , New\nYork: Oxford Umv. Press.\nGantt, W. H. 1938. Extension of a conflict based upon food to other\nphysiological systems and its reciprocal relaLions with sexual functions.\nAmer. J. Physiol„ 128, 78—74.\nGantt, W. II. 1944. Experimental basis for neurotic behavior: origin and\ndevelopment of artificially produced distmbances of behavior m dogs.\nPsychontn. Med. Monog., b, Nos. 3 and 4,\n1947. The genetics of schizophrenia. J. Abn. Soc. Psychol„\nGarrison, NT.\n42, 122-124.\nCassel, H. S. 1987. The control of excitation in the nervous system.\nHarvey Lect., pp. 169—198.\nGellerman, L. W. 1938. Form discrimination in chimpanzees and two-\nyear-old children: I. Form ( triangularity) per se. J. Genet. Psychol.,\n42\nGibbs. F. A. 1945. Electrical activity of the brain, Ann. Ret). Physiol.,\n1929. The reproduction of visually perceived forms.\nGibson, J• J.\nJ. Exp. Psychol., 12, 1-89.\n1941. A critical review of the concept af set in contempo-\nGibson, J• J.\nrary experimental psychology. Psychol. Bull., 98, 781—817.\nGibson, J. J. , and Crooks, L. E. 1938. A theoretical field-analysis of\nautdmobi!e dlivjng. Amer. J. Psychol., 51, 458—471.\nGillespie, W. H. 1944. The psychoneuroses. J. Ment. Sci., 90, 287—806.\nGoodenough, F. L, 1931. Anger in young children. Minneapolis: Univ.\nMinnesota Press.",
          "Receipt": {
            "TextSha256": "C997C360AD2B2CD6ED295882997C54D8EA10ED2BD4FD834D345438257E84C89B",
            "ProcessedAtUtc": "2026-09-16T23:48:46.7920093Z",
            "ImageSha256": "3A977DECF39D20DFCBD623EF650A6B1B138227CB16CB4FE4C8623EE75A18010A",
            "Language": "en-US",
            "Lines": 51,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000331.ocr.txt",
            "Image": "hebb-refs-000331.png"
          }
        },
        {
          "PdfPage": 332,
          "Text": "Bibliography\n809\nCreene, R., Paterson, A. S., and Pile, C. C S. 1945. Hypertrichosis with\nmental changes: the effect of adrenalectomy. Brit. Med. J., No. 4402,\n698-699.\nGuthrie, E. R. 1946. Psychological facts and psychological theory.\nPsychol. Bull. , 43, 1—20.\nHalstead, W. C. 1947. Brain and intelligence. Chicago: Univ. Chic.\nPress\nHanawalt, N. G. 1937. Memory trace for figures in 'recall and recogni-\nton. Arch. psychol , No. 216, 1—89.\nI-larns, H. J, 1944. Brucellows: a case report illushating a psychosomatic\nproblem. Psychosom. Med., 6, 3.34—335. ,\nHauptman, A. 1946 Capillaries in the finger nail fold in patients with\nneurosis, epilepsy, and rmgr$ne. Arch. Neurol. Psychiat., 56, 631—642.\nHead, H. 1920, Studies in neurology. London: Frowde, Hodder and\nStoughton.\nHeath, R. G. , and Pool, J. L. 1948. Bilatclal fractional resection of\nflontal coltcx for the treatment of psychoses. J. Nerv Ment. Dzs , 107\n411_429.\nHebb, D. O. 1937a. The innate organization of visual activity: I. Per-\nception of figures by rats reared in total darkness. J Ccnet. Psychol.,\n51, 101-126.\nHebb, D, O, 1937b. The innate orgamzation OE visual activity: II.\nTransfer of response in the discrimination of brightness and size by rats\nreared in total chukness. J. Comp. Psychol., 24, 277—299.\nHebb, D. O. 1938u. Studies of the m ganization of behavior: I. Behavior\nof the in a field orientation. J Comp. Psychol., 25, 383—352.\nHebb, D. O. 1988b. Studies of the orgamzation of behavior: II.\nChanges in the field orientation of the rat after cortical destruction.\nJ. comp. Psychot., 26, 427-444.\nHebb, D. O. 1989. Intelligence in man after large removals of cerebral\ntissue: report of four left frontal lobe cases. J, Cen. P5ffchol., 21, 78—87,\nHebb, D. O. 1942a. The effect of early and late brain injury upon test\nscores, and the nature of normal adult intelligence. Proc. Amer, Phil.\nsoc., 85, 275-292.\nHebb, D. O. 1942b. Verbal test material independent of special vocabu-\nlary difficulty. J. Educ. Psychol., 88, 691—696.\nHebb, D. O. 1945a. The forms and conditions of chimpanzee anger.\nBull. Canad. Psychol Ass., 5, 82—85.\nHebb, D. O. 1945b, Man's frontal lobes: A critical review. Arch.\nNeurol. Psychiat., 54, 10-24.\nIlebb, D. O. 1946a. Emotion in man and animal: an analysis of the\nintuitive processes OE recognition. Psüchot. Ret)., 53, 88—106.\nHebb, D. O. 1946b. On lhe nature of fear. Psychol. Ret., 58, 259—276.\nHebb, D. O. 1947. Spontaneous neurosis in chimpanzees: theoreti&l re-\nIn lions with clinical and experimental phenomena, Psychosom, Med„\n9, 3-16.",
          "Receipt": {
            "TextSha256": "614B18A368659899D22ABE1541371FA4D98CFBEF07369375B77CC66A94AC6B02",
            "ProcessedAtUtc": "2026-09-16T23:48:46.9806849Z",
            "ImageSha256": "96EFBD5C5B3A22A99B6F6B34A3BDE0C70DE913CE425440730302331CCCB74B7A",
            "Language": "en-US",
            "Lines": 48,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000332.ocr.txt",
            "Image": "hebb-refs-000332.png"
          }
        },
        {
          "PdfPage": 333,
          "Text": "810\nBibliography\nHebb, D. O. , and Foard, E. N. 1945. Errors of visual recognition and\nthe nature of the trace. J. Exp. Psychol., 85, 835—848.\nHebb, D. O. , and Morton, N. W. 1948. The McGill Adult Ccnnprehell„\nSion Examination: \"Verbal Situation\" and \"Plcture Anomaly\" Series.\n1. Educ. Psychol., 84, 16-25.\nHebb, D. 0., and Morton, N. W. 1944. Note on thc measmement of\nadult intelligence. J. Cen. Psychol., 80, 217—223.\nHebb, D. 0., and Penfield, W. 1940. Iluman behavior after extensive\nbilateral removal from the frontal lobes. Arch. Neurol Psychiat., 44,\n421-438.\nHebb, D. O. , and Riesen, A. H. 1948. The genesis of irrational [ears.\nBull Canad. Psychol. Ass., 8, 49—50.\nHebb, D. O. , and Williams, K. 194] Exoerimental control of cues de-\ntermining the rat's orientation. Bull. Canad. Psychol. Ass., 1, 22—28.\nHebb, D. 0., and Williams, K. 1946. A method of rating animal intel-\nligence. J. Cen. Psychot., 84, 59—65.\nHerrick, C. J. 1929. The thinking machine. Chicago: Univ. Chic. Press.\nHilgard, E. R., and Marquis, D. G. 1940. Conditioning and learning.\nNew York: Appleton-Century.\nHoagland, H 1947. Enzyme kinetics and the dynamics of behavior.\nJ. comp. Physiol. Psychol., 40, 107-127.\nHoagland, H., Malamud, W. , Kaufman, I. C.. and Pincus, a. 1946.\nChanges in the electroencephalogram and in the excretion of 17-keloster-\noids accompanying electro-shock therapy of agita Led depresslon. Psy-\nchosom. Med., 8, 246—251.\nHobbs, G. E. 1941. Mental disorder in one of a pair of identical twins.\nAmer. J. Psychiat., 98, 447450.\nHobhouse L. T. 1915. Mind in evolution. 2nd Ed. London: Mac-\nmillan.\nHoviand, C. 1.\n\"Inhibition of reinforcement\" and phenomena of\n1936.\nexperimental extinction. P? oc. Nat. Acad. Sci., Wash, 22, 480—438.\n(Quoted by Hilgard and Maaquis, 1940, p. )\nHovland, C. I. 1937. The generalization of conditloned responses. II.\nThe sensory generalization o} conditioned responses with varymg inten-\nsities of tone. J. Cenet. Psychol., 51, 279—291.\nHull, C. L. 19M. The concept of the habit-family hierarchy and mazo\nlearning. Psychol. Reu. , 41, 88—54; 184—152.\nHull, C. L. 1943. Principles of behavior: an introduction to bchacior\ntheory. New York: Appleton-Century.\nHull, C. L. 1945. The discrimination of stimulus configurations and the\nhypothesis of afferent neural interaction. Psychol. Rev., 52, 138—142,\nHumphrey, G. 1940. The problem OE the direction Of thought. Brit. J.\nPsychol., 80, 188-196.\nHunt, J. McV. 1941. The effects of infant feeding-frustration upon adult\nhoarding in the albino rat. J. Abn. Soc. Psychol., 86, 388—360.",
          "Receipt": {
            "TextSha256": "76F6935C54299BE5A61E5A8DE5EEC35CD063756B515AF798244960466863BE03",
            "ProcessedAtUtc": "2026-09-16T23:48:47.1812782Z",
            "ImageSha256": "B2F273276F2F8540EBFD9E5376411799E4B4B670C95FAD75D1A735E2B4B238F6",
            "Language": "en-US",
            "Lines": 49,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000333.ocr.txt",
            "Image": "hebb-refs-000333.png"
          }
        },
        {
          "PdfPage": 334,
          "Text": "Bibliography\n811\nHunter, W. S. 1984. Learning: TV. Experimental studies of learning.\nIn Murchison, C., Handbook of general experimental psychology.\nWorcester, Mass. : Clark Univ. Press, pp. 497—570.\nJackson, T. A. 1942, Use of the stick as a tool by young chimpanzees.\nJ. comp. Psycho!., 223-235.\nJacobsen, C. F., Jacobsen, M. M. , and Yoshioka, J. G. 19.32. Dei elop-\nmcnt of an infant chimpanzee during her first year. Comp. Psychol.\nMonog., 9, 1-94.\nJames, W. 1910. Principles of psychology. New York: Holt.\nJasper, H. H. 1987. Signs of beal activity. Psychol. Bull. ,\nJasper, H. H. 1941. Electroencephalography. In Penfield, W. , and\nErickson, T. C, Epilepsy and cerebral localization. Springfield: Thomas,\npp. 380-454.\nJasper, II. H., and Fortuyn, J. D. J 946. Experirnental studies on the\nfunctional anatomy of petit mal epilepsy. Publ. Ass. Res. Nerc. Ment.\nDis., 26, 272-298.\nJefferson, C. 1987. Removal of right or left frontal lobes in man. Brit.\nMed. J., 2, 199-206.\nJersild, A. T., and Holmes, F. B. 1985. Children's fears. New Yolk:\nTeach. coli. Btu. Publ,\nJolliffe, N. ] 942. The neuropsychiatric manifestations of vitamin defi-\nciencies, J. MI. Sinai liosv., 8, 658-667.\nJones, H. E. , and Conrad, H. S. 1933. The growth and decline of intel-\nligence: a study of a homogeneous group between the ages of ten and\nsixty. Genet. Psychol. Monog.z 18, No. 8.\nJones, H. E. , and Jones, M. C. 1928. A study of fear. Childhood Educ.,\n5, 136-148.\nJones, M. C. 1938. Emotional development. In Murchison, C., A hand-\nbook of child psychology. 2nd Ed. Worcester, Mass. : Chrk Univ. Press,\npp. 271-302.\nKappers, C. U. A., Huber, G. C, and Crosby, E. C. 1936. The compara-\ntioe anatomy Of the neroous system of vertebrates, including man. New\nYork: Macmillan, Vol. I.\nKarnosh, L. J. , and Gardner, W. J. 1941, An evaluation of the physical\nand mental capabilities following removal Of the righ& celebral hemi-\nsphere. Cleveland Qin. Quart., 8, 94—106.\nKennard, M. A. 1939. Alterations in response to visual stimuli following\nlesions of frontal lobe in monkeys. Arch. Net\" 01. Psychiat., 41, 1153—\n1165.\nKennard, M. A., and Ectors, L. 1938. Forced circling in monkeys follow-\ning lesions of the frontal lobes. J. Neurophysiot., 1, 45—54.\nKennedy, F. , and Wolf, A. 1936. The relationship of intellect to speech\ndefect in aphasic patients. J. Nerv. Ment. Dis., 84, 125—145; 293—811.\nKeschner, M„ Bender, M., and Strauss, I. 1988. Mental symptoms asso-",
          "Receipt": {
            "TextSha256": "D271DE2D66E9E31F3CB6A071F1A81F1CF6BB817BE09EBA04CB8149408E13C037",
            "ProcessedAtUtc": "2026-09-16T23:48:47.3924494Z",
            "ImageSha256": "E52B46EE82BD2C151CEDFB55DED61C10C01954D7519D8BA0FD271C4DBF4123E4",
            "Language": "en-US",
            "Lines": 46,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000334.ocr.txt",
            "Image": "hebb-refs-000334.png"
          }
        },
        {
          "PdfPage": 335,
          "Text": "812\nBibliography\nciated with brain tumou a study of 530 verified cases. J. Amer. Med.\nAgs., 110, 714-718.\neKinder, E. F. 1927. A study of the nest-building actisity of thc albino\nrat. J. zool., 47, 117-161.\nKinsey, A. C., Pomeroy, B. , and Martin, C. E. 1948. Sextla! behavior\nin the human anatc. Philadelphia: Saunders.\nKlebanoff, S. C. 1945. Psychological changes in organic brain lesions and\nablahons. Psychol. Bull, 42, 585—623.\nKleitman, N. 1939. Sleep and umkeftdness. Chicago: Univ Chic. mess.\nKlineberg, O. 1940. Sociat psychology. York' Holt.\nKoffka, K.\nKoffta, K.\nBracc.\nKöhler, W.\nKöhler, W.\nKohler, W.\nKöhler, W.,\n1924. The growth of the mind. New Yolk: Harcourt, Brace.\n1935. Principles of Gestalt psychology. New York: Harcourt,\n1925. The mentality of apes. New York: Harcourt, Brace.\n1929. Gestalt psychology. New York: Livelight.\n1940. Dynaniics in psychology. New York: Liveright.\nand Wallach, H. 1944. Figural after-effects: an investigation\nof visual processes. Proc. Amer. Phil Soc., 88, 269—857.\nKrechevsky, 1 1932. \"Hypotheses\" versus \"chance\" in the pre-solulion\nperiod in sensory discrimination-learning. Unit). Calif. Publ. Psychol ,\nKrechevsky, I. 1938. An experimental investigation of the principle of\nproximity in the visual perception of the rat. J. Exp. Psychol., 22,\n497-523.\nKubitschek, P. E. 1928. The symptomatology of tumols of the frontRl\nlobe based on a series of twenty-two cases. Arch. Neurol. Psychiut., 20,\n559-579.\nLandis, C. 1947. A modern dynamic psychology. J. Comp. Physiol.\nLandis, C and Hunt, W. A. 1932. Adrenalin and emotion. Psychol.\nLashley, K, S. 1029a. Brain mechanisms and intelligence: a quantitative\nstudy of iniuries to the brain. Chicago: Univ. Chic. Press.\nLashley, K. S. 1929b. Nervous mechanisms in learmng. In Murchi-\nson, C. , The foundations of experimental psychology. Worcester: Clark\nUniv. Press, LW, 524—568.\nLashley, K. S. 1980. Basic neural mechanisms in behavior. Psychol.\nReu, 87, 1-24.\nLashley, K. S. 1984. The mechanism of vision. VIII. The projection of\nthe retina upon the cerebral cortex of the rat. J. Comp. Neurol., 60,\n5M9.\nLashlgy, K. S.\nArch. Neurol.\nLashley, K. S.\nPsychot, Rev.,\n1987. Functional determinants of cerebral localization.\nPsychiat., 88, 871-887.\n1938a. Experimental analysis of instinctive behavior.\n45, 445—471.",
          "Receipt": {
            "TextSha256": "12C29121BD4A0FF21C0502EAEA00AEF57A1E9A3DAF35BB1E09DEE1047BB95013",
            "ProcessedAtUtc": "2026-09-16T23:48:47.565591Z",
            "ImageSha256": "DFC3888A0AF6F45A17AE27DA1136267356DE50C33704D2DA5E474AF528B7BFAB",
            "Language": "en-US",
            "Lines": 54,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000335.ocr.txt",
            "Image": "hebb-refs-000335.png"
          }
        },
        {
          "PdfPage": 336,
          "Text": "Bibliography\n818\nLashley, K. S. 1938b. The mechanism of vlsion: XV. Preliminary studies\nOF the mt's capacity for detail vision. J. Cen. Psychol., 18, 123—193.\nLashley, K. S. 193%. The thalamus and emotion. Psychol. Reu , 45,\n42—61.\nLashley, K. S. 1941. Patterns of cerebral integration indicated by the\nscotomas of migraine. Arch. Neural. Psychiat., 46, 831—839.\nLashley, K. S. 1942a. The problem of cerebral organization in vision.\nIn Kluver, H., Visual mechanisms. Biol. Sympos., 7, 301—822.\nLashley, K. S. 1942b. An examination of the \"continuity theory\" as ap-\nplied to learni:qg. J. Cen. Psychol,., 26, 241—265.\nLashley, K. S. 1944. Studies OE cerebral' function in XIII.\nApparent absence of trahscortical association in maze learning. J. Comp.\nNeurol., 80, 257-281.\nLashley, K. S. , and Clark, G. 1946. The cytoarchitecture of the cerebral\ncortex of Ateles: a critical examination of architectornc studies. J. Comp.\nNeurol., 85, 223-806.\nLashley, K. S. , and Wade, M. 1946. The Pavlovian theory of generahrza-\ntion, Psychol. Rev., 53, 72—87.\nLeeper, R. W. 1935. A study of a neglected portion of the field of\nleaming—the development of sensory organization. J. Genet. PsycJlDl., 46,\n41-75.\nLeeper, R. W. 1948. A motivational theory of emotion to replace \"Emo-\ntian as disorganized response.\" Psychol. Reu. , 55, 5—21.\nLehmann, J. E, 19870. Tlle effect of changes in the potassium-calcuun\nbalance on the action of mammalian A nerve fibers. Amer. J. Physiol..\n118, 613-619.\nLehmann, J. E. 1937b. The effect of changes in pH on the action of\nmammalian A nerve fibers. Amer. J. Physiol., 118, 600—612.\nLevine, j. 1945m Studies in the interrelations of central nervous struc-\ntures in binocular visrjn: I. The lack of bilateral bansfe• of visual dis-\ncriminative' habits acquired monocularly by the pigeon. J. Genet. P.S!,'-\nchot., 67, 105-129.\nLevine, J. 1945b. Studies in the interrelations of central nervous Struc-\ntures in binocular vision: II. The conditions under which interocular\nfransfer of discriminative habits takes place in the pigeon. J. Genet.\nPsychol., 67, 131-142.\nLewin, K. 1938. Will and needs. In Ellis, W. D., A source book of\nGestalt psychology. London: Kegan Paul, Trench, Tmbner, pp. 283—299,\nLibet, B., and Gerard, R. W. 1939. Control of the potential rhythm of\nthe isolated frog brain, J. Neurophysiol., 2, 153—169.\nLiddell, H. S. 1938. The experimental neurosis and the problem of\nmental disorder. Amer. J. Psychiat., 94, 1035—1041.\nLiddell, H. S. 1944. Animal bchavior studies bearing on the problem of\npain. Psychosom. Med., 6, 261--268.\nLorente de N6, R. 1938\". Synaptic stimulation of motoneurons as a local\nprocess. J, NeurophysioZ., I, 195—206.",
          "Receipt": {
            "TextSha256": "F3013B7D72FBA9867762E414941D37C92C80B0A152D7B618B2E7E09CBB7A3D27",
            "ProcessedAtUtc": "2026-09-16T23:48:47.7584734Z",
            "ImageSha256": "131BCADFE7E28DB8524D2033C4326C543B6367573CEF2F79115FBABD464F2144",
            "Language": "en-US",
            "Lines": 48,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000336.ocr.txt",
            "Image": "hebb-refs-000336.png"
          }
        },
        {
          "PdfPage": 337,
          "Text": "Bibliography\nLorente de N6, H. 1988b. Analysis of lhe acuvity of the chains of inter-\nnuncial neurons. J. Neurophysiol , 1, 207—244.\nLorente de NO, R. 1989. Transmission of impulses through cranial motor\nnuclei. J. Neurophysiot., 2, 402—464.\nLorente de N6, R. 1948. Cerebral cortex: architecture. In Fulton, J.\nPhysiology of the nervous system. 2nd Ed. New York. Oxford Univ.\nPress, pp. 274-801.\nLorenz, K. 1985. Del Kumpan in der Umwelt des Vogels. J. Ornith., 8b,\n137-218; 289-413.\nLoucks, R. B. 1935. Experimental delimitation of neural structures essen-\ntiai for learning: The attempt to condition striped muscle responses With\nfaradization of the sigmoid gyli. J. Psychol., 1, 5—44.\nLoucks, R. B. 1938. Studies of neural Sh•uctures essential for learning.\nII. The conchtioning of salivary and striped muscle responses to faradl-\nzation of cortical sensory elements, and the acllon of sleep upon such\nmechanisms, J. Comp, Psychol., 25, 315—882.\nLuekhardt, A. B., and Carlson, A. J. 1915. Contributions to the physi-\nology of the stomach. X VIL On the chemical control of the gastric\nhunger mechanism. Amer. J. Physiol., 36, 37—46.\nMcBnde, A. F., and Hebb, D. O. 1948. Behavlor of the captive bottle-\nnose dolphin, Tursiops truncatus. J. Comp. Physiol. Psychol., 41, Ill—\n128.\nMcCulloch, T, L., and Haslerud, G. M. 1939. Affective responses of an\ninfant chimpanzee reared in isolation from its kind. J. Comp. Psychol.,\n28, 437—445.\nMcCulloch, W. S. 1944u. Cortico-cortical connections. In Buoy, P., The\nmecentral motor cortex, Urbana, Ill.: Univ. Illinois Press, pp. 218—242.\nMcCulloah, W. S. 1944b. The functional uganization of the cerebral\ncortex. Physiol. Ret'., 24, 390—407.\nMcGeoc•h, J. A 1942. The psychology of human learning. New York:\nLongmans, Green.\nMaier, N. R. F. , and Schneilla, T. C. 1935. Principles of animal psy-\nchology. New York: McGraw-Hill.\nM',ushall, W. H., and Talbot, S. A. 1942. Recent evidence for nemal\nmechanisms in vision leading to a general theory of sensory acuity. In\nKlüver, H., Visual mechanisms. Biol. Symp., 7, 117-164.\nMasserman, J. H. 1942. The hypothalamus in psychiatry. Amer. J.\nPsychiat., 98, 638-687.\nMasserman, J. H. 1943. Behavior and neurosis• An experimental psycho-\nanalytic approach to psychobiologic 73Tinciples. Chicago: Univ. Chic.\nPress.\nMatthews, R. S. 1938. Pellagra and nicotinic acid. J. Amer. Med. Ass.,\nIN, 1148-1158.\nMiller, G. A. 1947. The masking of speech. Psychol. Bull, 44, 105—129,\nMiner, J. B. 1905. A case of vision acquired in adult life. Psychol. Ret).\nMonog. suppl, 6, No. 5, 103-118.",
          "Receipt": {
            "TextSha256": "B04C77A3F1361C732964C135B396D600D715FAD377BF3829E53360DA6F5A2F57",
            "ProcessedAtUtc": "2026-09-16T23:48:47.9394556Z",
            "ImageSha256": "3706D6BED490FB467A3888E4F8BA2AE1B1A6DAEAC2216351D19584A85C501711",
            "Language": "en-US",
            "Lines": 47,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000337.ocr.txt",
            "Image": "hebb-refs-000337.png"
          }
        },
        {
          "PdfPage": 338,
          "Text": "Bibliography\n315\nMixter, W. J., Tillotson, K. J. , and Wies, D. 1941. Reports of partial\nfrontal lobectomy and frontal lobotomy performed on three patients: one\nchronic epileptic and two cases of chronic agitated depression. Psy-\nchosom. Med. , 8, 26—37.\nMOI gan, C. T. 194B. Physiological psychology. New York; McCmw-Hil\\.\nMonson, R. S., and Dempsey, E. W. 1948. Mechanism of thalamocortical\naugmentation and repetltion. Amer. J. Physiol., 188, 297—308.\nMoss, F. A. ( Ed.). 1942. Comptnatice psychology. Rec•. Ed. New York:\nPrentice-Hal].\nMowrer, O. H. 1941. Motiva\"ion and learning in relation to the national\nemergency. Psychol. Bull. , 38, 421—481.\nMowrer, O. 11., and Mowrer, W. M. 1988. Enuresis—a method for its\nstudy and treatment. Amer. Orthopsychiat., 8, 436—459,\nMuenzinger, K. F. 1934. Motivation in lemming. I. Electric shock for\ncorrect response in the visual discrimination habit. J. Comp. PsychoZ.,\n17, 267-277.\nMurphy, J. P., and Gellhorn, E. 1945. Further investigations on dien-\ncephalic-cortical relations and their significance for the problem cf emo-\nlion. J, Neurophysiol., 8, 431—447.\nNafe, J. P. 1984. The pressure, pain, and temperature senses. In Murchi-\nson, C., Handbook of general experimental psychology. Worcester,\nMass.: Clark Univ. Press, pp. 1037-1087.\nNauta, W. J. H. 1946. Hypothalamic regulation of sleep in rats: An ex-\nperimental study. J. Neurophysiol„ 9, 285416.\nNeet, C. C. 1983. Visual pattern discrimination in the Macacus rhesus\nmonkey. J. Genet. Psychol., 43, 168—196.\nNeff, Walter S. 1938. Socioeconomic status and intelligence: a critical\nsurvey. Psychot. Bull, 85, 727-757.\nNichols, 1., and Hunt, J. McV. 1940. A case of partial bilateral frontal\nlobectomy: a psychopathological study. Amer. J. PsycJtiat, 96, 1068—\n1083.\nNissen, H. W., Machover, S., and Kinder, E. F. 1985. A study of per-\nfonnance tests given to a group of native African negro children, Brit.\nJ. Psychol, 25, 308-355.\nPavlov, l. P. 1927. Conditioned reflexes. Oxford: Humphrey Mlford.\nPavlov, l. P. 1928. Lectures on conditioned reflexes. New York: Inter-\nnational.\nPavlov, l. P. 1932. The reply of a physiologist to psychologists. Psychol.\nReu, 89, 91-126.\nPenfield, W., and Jasper, I-I. 1946. Highest level seizures. Rcs. Publ.\nAss. New. Ment. Dis., 26, 252-271.\nPennington, L. A. 1938. The function of the brain in auditory Ituiza-\ntion. I V. Method of training and control experiments. J. Contp\nPsychol., 25, 195-211.\nPilgrim, F. J., and Patton, R. A. 1947. Patterns of self-selection of puri-",
          "Receipt": {
            "TextSha256": "F467D18D78F16EC7043B10514C3804859438A17706388FF4D9A1F3CC7B4B7797",
            "ProcessedAtUtc": "2026-09-16T23:48:48.1248991Z",
            "ImageSha256": "9AA7A943CA1B68209100827B68B3F1AD3AE977C1610F58ED7197F8E17E68F4AE",
            "Language": "en-US",
            "Lines": 47,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000338.ocr.txt",
            "Image": "hebb-refs-000338.png"
          }
        },
        {
          "PdfPage": 339,
          "Text": "316\nBibliography\nJ. Comp. Physiol Psychol., 40,\nfied dietary components by the rat.\nPillsbury, W. B. 1918. \"Fluctuations of attention\" and the refractory\nperiod. J. Phil. Psychol. Sci. Nleth., 10, 181—185.\nPolyak, S. L. 1941. The 7 etina. Chicago: Univ. Chic. P.ress.\nPostman, L. 1947. The and present status of the law of effect.\nPsychol. Bull. , 44, 489—568.\nPratt, C. C. 1939. The logic of modern psychology. New York: Mac-\nmillan.\nPrentice, W. C. H. 1946. Operationism and psychological theory: a note.\nPsychol. Rev., 53, 247-24g.\nProsser, C. L. 1934. Action potentials in the nervous system of the cray-\nfish: I. Spontaneous impulses. J. Cell. Comp. Physiol, 4, 185—209.\nRichter, C. P. , Holt, L. E., and Barelare, B. 1988. Nutritional require-\nments for normal growth and reproduction in rats studied by the self-\nselection method. Amer. J. Physiol., 122, 784—744.\nRlddoeh, C. 1941. Phantom limbs and body shape. Brain, 64, 197—222.\nRiesen, A. H. 1947. The development of visual perception in man and\nchimpanzee. Science, 106, 107—108.\nRiesen, A. H., and Nissen, H. W, 1942. Non-spatial delayed response by\nthe matching technique. J. Comp. Psychol., 84, 307—818.\nRoethlisburger, F. J., and Dickson, W. J.\n1939. Management and the\nteorker. Cambridge: Harvard Univ. Press.\nRomano, J., and Coon, C P. 1942. Physiologic and psychologic studies\nin spontaneous hypoglycemia. Psychosom. Med., 4, 288—300.\nRowe, S. N. 1937. Mental changes following the removal of the right\ncerebral hemisphere for brain tumor. Amer. J. Psychiat., 94, 605—614.\nRubin, 1921. Visuell wahrgenommene Figuren: Studien in psycho-\nlogischcr Analyse. Teil I. Berlin: Gyldendalske Boghandel.\nRylander, G. *1989. Personality changes after operations on the frontal\nlobes: a clinical study of 82 cases. London: Humphrey Milford.\nSaphir, W. 1945. Chronic hypochloremia simulating psychoneurosis.\nJ. Amer. Med. Ass., 129, 510-512.\nSchneirla, T. C. 1948. Psychology, comparative. Encycl. Brit.\nScott, W. W., Scott, C. C. , and Luckhardt, A. B. 1988. Observations on\nthe blood sugar level before, during, and after hunger periods in humans.\nAmer. J. Physioi., 198, 248-247,\nSenden, M. v. 1932. Raum- und Gestaltauffassung bei operierten Blind-\ngeborenen 007' und nach der Operation. Leipzig: Barth.\nSherrington, C. S. 1903. Integrat!ce action of the nercous system. New\nYork: Scribner.\nSherrington, C. S. 1925. Remarks on some aspects of reflex inhibition.\nPec. Rou. soc., 97B, 519-545.\nSherrington, C. S. 1941. Man on his nature. New York: Macmillan.\nSkinner, B. F. 1988. The behavior of organisms: an experimental analysis.\nNew York: Appleton-Century.",
          "Receipt": {
            "TextSha256": "98B59445136D5AFE42BD8E12822FE9F8A85B0CB846927084F77FEC7BD97AF733",
            "ProcessedAtUtc": "2026-09-16T23:48:48.3228294Z",
            "ImageSha256": "EF8EA9D5553FB198390F6EEC4CE0058C0620F48E6A3FD0ED9F8DB187D1978876",
            "Language": "en-US",
            "Lines": 49,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000339.ocr.txt",
            "Image": "hebb-refs-000339.png"
          }
        },
        {
          "PdfPage": 340,
          "Text": "Bibliography\n317\nSmith, D. E. 1939. Cerebral localization in somesthetic discriminatlon in\nthe rat. J. comp. Psychol., 28, 161-188.\nSmith, K. U. 19.36. Visual discrirmnation in the cat: III. The 1 elative\neflect OE paired and unpaired st1muLi in the d:scrmainative behauor of\nthe cat. J. Genet. Psychol., 48, 29—57.\nSmith, W. K. 1945. The functional signlficance of the rostral cingular\ngyrus as revealed by its responses to electrical exci!ation. J. Neuro-\nphysiol., 8, 241-255.\nSpearman, C. 1927. The abilities of man. New York: Macmillan.\nSpence, K. W. 1938. Graduxl versus sudden solution of discrimination\nproblems by chimpanzees. J. Comp. Psychbl., 25, 213—224.\nSpence, K. W. 1940. Continuous versus non-continuous interpretations of\ndiscrimination learning. Psychol. Ret. , 47, 271—288.\nSpeny, R W. 1948. Visuomotor cool dination in the newt ( Triturus\nviridescens) after regeneration of the ophc nerve. J. Comp. Neurol., 79,\nSperry, R. W. 1947. Effect of crossing nerves to antagonistic limb mus-\ncles in the monkey Arch. Neurol. Psychiat., 58, 452—473.\nSpiegel, E. A , Millen, H. R. , and Oppenheimer, M. J. 1940. Forcbrain\nand rage reactions. J. N ctuaphvsioL, 3, 539—548.\nspies, T. D., Aring, C. D., Gelperin, J., and Bean, W. B. 1988. The\nmental symptoms of pellagra: Their reheE with nicotinic acid. Anwr. J.\nMed. sci„ 196, 461—475.\n1940. Morphine addiction in chimpanzees. Comp.\nSpragg, S. D S.\nPsychol. Monog., 15, No. '7.\nStoddard, C. D. , and Wellman, B. L. 1940. Environment and the IQ.\nYearb. Nat. soc. stud. Educ., 89 405-442.\nStookey, B., Scarff, J. , and Teitelbaum, M. 1941. Frontal lobettomy in\nthe treatment of brain tumors. Ann. Surg., 118, 161—169.\nSwank, R. L. , and Marc%lnd, W. E. 1946. Combat neuroses: develop-\nment of combat exhaustion. Arch. Neurol. Psvchiet., 55, 286—247.\nThorndikc, E L. 1931. Human learning. New York: Century.\nL. 1935. The vectors of mind. Chicago: Univ. Chic. Press.\nTinbergen, N. 1942. An objectivistic study of the innate behavior of\nanimals. Bibl. Biotheoret., Leiden, 1, 89—98.\nTinklepaugh, O. L. 1928. An experimental study of revresentative fac-\nt01S in monkeys. J. Comp. Psychol., 8, 197—236.\nTitchener, E. B. 1920, Notes from the psychological laboratory of Cornell\nUniversity. Amer. J. Psychol. 81, 212—214.\nTolman, E. C. 1982. Purposive behavior in animals and men. New\nYork: Century.\nTryon, R. C, 1989. Studies in individual differences in maze learning: VI.\nDisproof of sensory components: experimental effects of stimulus 'Liria-\n{ion. J. Comp. 28, 861-415.\nValentine, C. W. 1930. The innate bases of fear. J. Cenct. Psychol., 97,\n94—419.",
          "Receipt": {
            "TextSha256": "1454226F3EB510EA698E6479F6FAF9B2677D06F4311AE4E0C5F036619CF8D086",
            "ProcessedAtUtc": "2026-09-16T23:48:48.5131808Z",
            "ImageSha256": "25A970A8D1B36330896BF01E00BC05925E1C6492664F6B1718933F1AB2F9B19E",
            "Language": "en-US",
            "Lines": 48,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000340.ocr.txt",
            "Image": "hebb-refs-000340.png"
          }
        },
        {
          "PdfPage": 341,
          "Text": "818\nBibliography\nWalker, A. E., and Weaver, T. A. 1940. Ocular movements hom the\noccipital lobe in the monkey. J. Neurophysiol., 8, 858—857.\nWatson, J. B. 1924. Behaviorism. New York: Ion.\nWatts, J. W., and Freeman, W. 1946. Psychosurgery for the relief of in-\ntractable pain. J. Int. Colt. Surg., 9, 679—683.\nWechsler, D. 1989. The measurement of adult intelligence. Baltimore:\nWilliams and Wilkins.\nWeddel, C., Sinciair, D. C. , and Feindel, W. I-I. 1948 An anatomical\nbasis for alterations in quality OF pain sensibility. J. Neurophysiol., 11,\nWeisenbutg, T. , and McBride, K. E. 1935. Aphasia: a clinical and psy-\nchological study. New York: Commonwealth Fund.\nWeisenburg, T., Roe, A., and McBride, K E. 1986. Adult intelligence:\na psychological study of test performances. New York: Commonwealth\nFund.\nWeiss, P. 1941a. Autonomous versus reflexogenous activity of the central\nnervous system. Proc. Amer. Phil. Soc., 84, 58—64.\nWeiss, P. 1941b. Nerve patterns: The mechanics of nervc growth.\nGrowth ( Third Growth Symposium), 5, 168—203.\nWerner, H., and Strauss, A. 1989. Types of visuo-motor activity in their\nrelation to low and high performance ages. Proc. Amer. Ass. Ment.\nDefic., 44, 168-168.\nWilson, G., and Rupp, C. 1947. Present trends in thc practice of neu-\nrology. J. Amer. Med. Ass., 188, 509—511.\nWolf, E , and Zerrahn-Wolf, G. 1987. Flicker and the reactions of bees\nto flowers, J. Cen, Physiol., 20, 511—518.\nWolf, G. A., and Wolff, H. C. 1946. Studies on the nature of certain\nsympbms associated with cardlovascular disorders. Psychosom. Med, 8,\n293-319.\nWolff, H. C. e1948. Emotions and gastric function. Science, 98, 481—484.\nWolff, H. C., and Hardy, J. D. 1947. On the nature of pain. Physiol.\nRet,., 27, 167-199.\nWoodrow, H. 1927. The effect Of type of training on transference.\nJ. Educ. Psychol, 18, 160-171.\nWoodworth, R. S. 1921. Psychology. New York: Holt.\nWoodworth, R' S. 1988, Experimental psychology. New York: Holt.\nWortis, H. , Stein, M. H. , and Jolliffe, N. 1942. Fiber dissociation in\nperipheral neuropathy. Arch. Int. Med., 69, 222—237.\nYerkes, R. M. 1916. The mental life of monkeys and apes: a study Of\nideational behavior. Behavior Monog., 8, No. 1.\nYoung, P. T. 1941. The experimental analysis of appetite. Psychol. Bull,\n129-164.\nYoung, P. T. 1944. Studies of food preference, appetite and dietary habit.\nl. Running activity and dietmy habit of the rat in relation to food pref-\nerence. J. Comp. Psychol., 87, 827—370.",
          "Receipt": {
            "TextSha256": "740091B73A88D5D964052C43525ED32992850C6B63C2D1CA4A26F22BC45749D0",
            "ProcessedAtUtc": "2026-09-16T23:48:48.6950145Z",
            "ImageSha256": "9E852041F40B3B48F0EEB99BB15AA0F4A45EEC91C236B86A7C5BE7046CB6AEF8",
            "Language": "en-US",
            "Lines": 46,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000341.ocr.txt",
            "Image": "hebb-refs-000341.png"
          }
        },
        {
          "PdfPage": 342,
          "Text": "Bibliography\n819\nZangwill, O. L. 1937. A study of the slgnificanc•c of attitude in lecogni-\nLion. Brit, J. Psychat., 28, 12-17.\nZener, K. 1,937 The significance of behavior accompanying conditioned\nsalivary secretion fol theones of the conditioned response.\n-Aincr. J.\nZollingel, R. 1935. Remmal left cerebral hemispheu_.'. tepui t of a case\nArch. Ncurot. Psychiat., 84, 1055—1064.",
          "Receipt": {
            "TextSha256": "16B0F676040B8D841D31D07C28E1C988F4AF714575FD04F636865410B5455FAF",
            "ProcessedAtUtc": "2026-09-16T23:48:48.7610067Z",
            "ImageSha256": "61726F30E91ED791B630A3081C18854E2CFF5AB13036EE3CCC919A35D2916C73",
            "Language": "en-US",
            "Lines": 9,
            "Engine": "Windows.Media.Ocr",
            "Output": "hebb-refs-000342.ocr.txt",
            "Image": "hebb-refs-000342.png"
          }
        }
      ]
    },
    {
      "Slug": "turing",
      "Paper": "Computing Machinery and Intelligence",
      "AtlasYear": 1950,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.cs.toronto.edu/~frank/csc2501/Readings/R1_Turing/Turing-1950.pdf",
      "PdfSha256": "F2D8540600CA8A94897CE7EAF35D99D336D4B39191FEE4481696487C2369B3B4",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 573,
          "EndLine": 585,
          "PdfPages": [
            29
          ],
          "Text": "Samuel Butler, Erevhon, London, 1865. Chapters 23, 24, 25, T h e Book of the ,IIachines.\nAlonzo Church, \" An Unsolvable Problem of Elementary Number Theory \",\nAmerican J . of ,?lath., 58 (1936), 345-363. K. Godel, \" fiber formal unentbcheidbare Satze der Principla Rlathematica\nund x-erwandter Systeme, I \", -Ilotzatshefte fur Math. u?zd Phys.,\n119311, 173-189. D. R . Hartree, Cnlculating Instrztments and ~IIachines,New York, 1949.\nS. C. Kleene, \" General Recursive Functions of Xatural Numbers \",\nAmerican J . of Math., 57 (1935), 153-173 and 219-244.\nG. Jefferson, \" The Mind of Mechanical Man \". Lister Oration for 1949.\nBritish ,Ilerlical Journal, vol. i (1949), 1105-1121. Countess of Lovelace, ' Translator's notes t c an artlcle on Babbage's\nAnalytical Engire ', Scient~fic ,Ilemozrs (ed. by R. Taylor), vol. 3\n(1842), 691-731. Bertrand Russell, History of Western Philosophy, London, 1940. A. 31. Turing, \" On Computable Xumbers, with an Application to the\nEntscheidungsproblem \", Proc. London ..lIath. Soc. (2), 42 (1937),\n230-265."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "logic-theory-machine",
      "Paper": "The Logic Theory Machine",
      "AtlasYear": 1956,
      "Status": "indexed",
      "Method": "windows-ocr-bibliographical-footnotes",
      "SourceUrl": "https://www.rand.org/content/dam/rand/pubs/papers/2024/P868.pdf",
      "PdfSha256": "FDC2CE8D519264B3B4F2D29F0F7B86A849E715ACB5C4F21677AB5F07B9D85F55",
      "Sections": [
        {
          "Section": "Bibliographical footnotes",
          "PdfPages": [
            5,
            10
          ],
          "Text": "Footnote 2 (printed page 3): B. V. Bowden (ed.), Faster Than Thought (London: Pitman, 1953), pp. 181-198.\nFootnote 5 (printed page 8): A. N. Whitehead and Bertrand Russell, Principia Mathematica, vol. I, 2nd edition (Cambridge: 1925).\nFootnote 5 (printed page 8): D. Hilbert and W. Ackermann, Principles of Mathematical Logic (New York: Chelsea, 1950), Chapter I."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "The 64-page RAND P-868 scan was processed with OCR. This report has bibliographical footnotes rather than a terminal bibliography. The three external works in footnotes 2 and 5 are indexed, with the source images checked; internal references, acknowledgments and report-date statements are excluded."
      ],
      "OcrPageCount": 64,
      "BibliographyEdition": "RAND P-868, 12 July 1956",
      "ReferenceCount": 3
    },
    {
      "Slug": "perceptron",
      "Paper": "The perceptron: A probabilistic model for information storage and organization in the brain",
      "AtlasYear": 1958,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://web.stanford.edu/class/psych209a/ReadingsByDate/01_30/Rosenblatt58Perceptron.pdf",
      "PdfSha256": "62B4F2C8C98719C76C7279FB708D3670894D9DE65488D1764E0CE80A59DE7104",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 543,
          "EndLine": 561,
          "PdfPages": [
            23
          ],
          "Text": "1. ASHBY, W. R. Design for a brain. New York: Wiley, 1952.\n2. CULBERTSON, J. T. Consciousness and behavior. Dubuque, Iowa: Wm. C. Brown, 1950.\n3. CULBERTSON, J. T. Some uneconomical robots. In C. E. Shannon & J. McCarthy (Eds.), Automata studies. Princeton: Princeton Univer. Press, 1956. Pp. 99-116.\n4. ECCLES, J. C. The neurophysiological basis of mind. Oxford: Clarendon, 1953.\n5. GOLDSTEIN, K. Human nature in the light of psychopathology. Cambridge: Harvard Univer. Press, 1940.\n6. HAYEK, F. A. The sensory order. Chicago: Univer. Chicago Press, 1952.\n7. HEBB, D. O. The organization of behavior. New York: Wiley, 1949.\n8. KLEENE, S. C. Representation of events in nerve nets and finite automata. In C. E. Shannon & J. McCarthy (Eds.), Automata studies. Princeton: Princeton Univer. Press, 1956. Pp. 3-41.\n9. KOHLER, W. Relational determination in perception. In L. A. Jeffress (Ed.), Cerebral mechanisms in behavior. New York: Wiley, 1951. Pp. 200-243.\n10. McCuLLOCH, W. S. Why the mind is in the head. In L. A. Jeffress (Ed.), Cerebral mechanisms in behavior. New York: Wiley, 1951. Pp. 42-111.\n\n11. MCCULLOCH, W. S., & PITTS, W. A logical calculus of the ideas immanent in nervous activity. Butt. math. Biophysics, 1943, S, 115-133.\n12. MILNER, P. M. The cell assembly: Mark II. Psychol. Rev., 1957,64,242252.\n13. MINSKY, M. L. Some universal elements for finite automata. In C. E. Shannon & J. McCarthy (Eds.), Automata studies. Princeton: Princeton Univer. Press, 1956. Pp. 117-128.\n14. RASHEVSKY, N. Mathematical biophysics. Chicago: Univer. Chicago Press, 1938.\n15. ROSENBLATT, F. The perceptron: A theory of statistical separability in cognitive systems. Buffalo: Cornell Aeronautical Laboratory, Inc. Rep. No. VG-1196-G-1, 1958.\n16. UTTLEY, A. M. Conditional probability machines and conditioned reflexes. In C. E. Shannon & J. McCarthy (Eds.), Automata studies. Princeton: Princeton Univer. Press, 1956. Pp. 253-275.\n17. VON NEUMANN, J. The general and logical theory of automata. In L. A. Jeffress (Ed.), Cerebral mechanisms in behavior. New York: Wiley, 1951. Pp. 1-41.\n18. VON NEUMANN, J. Probabilistic logics and the synthesis of reliable organisms from unreliable components. In C. E. Shannon & J. McCarthy (Eds.), Automata studies. Princeton: Princeton Univer. Press, 1956. Pp. 43-98."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "lisp",
      "Paper": "Recursive Functions of Symbolic Expressions and Their Computation by Machine, Part I",
      "AtlasYear": 1960,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www-formal.stanford.edu/jmc/recursive.pdf",
      "PdfSha256": "3D981849E59505EFF3F14397A177B409F5D978D43D114BDD67C956E74320FC92",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 554,
          "EndLine": 558,
          "PdfPages": [
            34
          ],
          "Text": "1. J. McCARTHY, Programs with common sense, Paper presented at the Symposium on the Mechanization of Thought Processes, National Physical Laboratory, Teddington, England, Nov. 24-27, 1958. (Published in Proceedings of the Symposium by H. M. Stationery Oﬃce).\n2. A. NEWELL AND J. C. SHAW, Programming the logic theory machine, Proc. Western Joint Computer Conference, Feb. 1957.\n3. A. CHURCH, The Calculi of Lambda-Conversion (Princeton University Press, Princeton, N. J., 1941).\n4. FORTRAN Programmer’s Reference Manual, IBM Corporation, New York, Oct. 15, 1956.\n5. A. J. PERLIS AND K. SAMELS0N, International algebraic language, Preliminary Report, Comm. Assoc. Comp. Mach., Dec. 1958."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "sir",
      "Paper": "SIR: A Computer Program for Semantic Information Retrieval",
      "AtlasYear": 1964,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://bitsavers.trailing-edge.com/pdf/mit/ai/aim/AITR-220.pdf",
      "PdfSha256": "62B359D2BDB995070BE4BC1473936A4CBFA33C61F8A68501BE11FF496D4B5171",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 6434,
          "EndLine": 6615,
          "PdfPages": [
            143,
            144,
            145
          ],
          "Text": "1. ACF Industries, Avion D'-v* tTr6nslating From Ordinary Discourse Into ormal Logic ii- A reliminary Stu4y#14 Scientific Report AF CRC,-TN-56-770.\n\n2s Bennett, J. Lo \"A Computer Program for Word Relations,,\" Memo 1961-1, Mechanical Translation Group, RLE, MIT. Cambridge) Mass. 1961*\n\n3. Bobrow D G. \"'Syntactic Analysis of English by Computer -- A Survey,\" Proc. FJCC# Spartan Press) 1963.\n\n4. lobrow, D. G., and Raphael, B. \"A Comparison of List-Processing Computer Languages,\" Cqo_m\"m. ACM# May or June 1964.\n\n5. Carnap, R. Meani Illinois. 947.\n\nand Neces\n\nU. of Chicago Press, Chicago,\n\n6. Carnap, R. \"Foundations of Logic and Mathematics,\" International Encyclo'zedia of Unified Science, Volume 1, no. 3 U. Of Chicago Press, Chicago, Illinois. 939.\n\n7. Carroll, J. Do, Abelson, R P., and Reinfeld, W. \"A Computer Program Which Assesses the Credibility of Assertions,\" draft. Yale University. July 1963.\n\n8. Charney, E. \"Word-meaning and Sentence-meaning,\" abstract in\n\nMechanical Translation, Volume 7 no- 2. 1963.\n\n9. Chomsky, N.\n\nic Structu\n\nMouton and Co. 1947.\n\n.Xntact\n\nres\n\n\"Picture Processing in a Picture Language Machine,\"\n\n10* Cohen, D.\n\nNational Bureau of Standards Report 7885* April 1962.\n\n11. Corbato, F. J., et. al. The CoMEatible Time\"Sharina Systemo MIT Press, Cambridge, Mass. 1963.\n\n12. Darlington, J. L. \"Translating Ordinary Language into Symbolic Logic,\" abstract in Mechanical Translation, Volume 7 no. 2. 1963*\n\n13. Davis, M., and Putnam, H. \"A Computational Proof Procedure,\" AFOSR TR 59-124. Rensselaer Polytechnic Institute, Troy, N#Y. 1959.\n\n14. Feigenbaum, E. \"The Smulation of Verbal Learning Behavior,\" Proco WJCC, Volume 19. 1961.\n\n15. Freudenthal, H. LINCOS,* Design of a Lan&uaaefor Cosmic Intercourse. North Holland Press, 1960.\n\n\f144\n\n16. Fries. 1952.\n\nThe Structure ofEng_lisho Harcourt, Bracep New York,\n\n17* Green, Be Fe, Jr*, etft ale \"Baseball: An Automatic QuestionAnswerer,\" Proc, WJCC, Volume 19. 1961.\n\n18. Kazemier, B. H., and Vuysje, De, eds. The Concept-a'd the Role of the Model in Mathematics and Natural and Social Sciences# Gordon and Breach Science Publishers, N.Y. 1963.\n\n19. Klein, S. \"Some Experiments Performed with an Automatic Para-\n\n.pffiraser,\" 1963.\n\nabstract in Mechanical Translation, Volume 7\n\nno* 2,\n\n20. Kochen, Ma \"Experimetital Study of 'Hypothesis Formation' by Computer,\" Proc. 4th London mEosium on Information Th C. Cherry, ed, London, 1961.\n\n21. Lindsay, R. K.\n\nA Program for Parsing Sentences and Making\n\nInferences about Kinship Relations,\" Proc'&Western Management\n\nScience Conference on Simulation, A. Hoggatt, ede, to be\n\npublished.\n\n22. McCarthy, Jo \"Programs with Common Sense,\" ProcII-Symposiu on Mechanization of Thou&ht Processeso National Physics Laboratory, Teddington, England. Her Majesty's Stationery Office, London. 1959,\n\n23. McCarthy, Jo, et. al. LISP 1 5 Prorammerls Manual. MIT Press, Cambridge, Mass* 1963.\n\n24. Maron, I. RAND Corp., Santa Monica, California. private communication, 1963.\n\n25a Minsky, M# \"Steps To-ward Artificial Intelligence#\" special computer issue, 1961.\n\nProco, IRE,\n\n26. Newell, A., edo Information Proces Prentice Halli Englewood Cliffs, N.\n\nLapauage V Manual. 1961.\n\n27a Newell, A., et, al. \"Empirical Explorations of the Logic Theory\n\nMachine A Case Study in Heuristics,\" Proc. WJCC,\n\n1957a\n\n28. Newell, A., et. al. \"Report on a General Problem-Solving Program,\"\n\nProc. International Conference on Iformation Processi\n\nParis,\n\nUNESCO House, 1959.\n\n29. O'Donnell, M, The New D2,y.In and Day Out. Row, Peterson and Co. Evanston, Illinois, 1948.\n30. Oden C K, Basic_English, Paul, Trench, Trubner and Co. London. 1 32 a\n\nI\n\n\f145\n\n31. Phillips$ A. V. \"A Question-oAnswering Routine,\" MIT Mathematics Dept. Mastt,--r's Thesis. Cambridge, Mass. 1960,\n\n32. Quillian, R.\n\nRevised Desi n for an Understanding Machine\n\nMechanical Translatl\n\nVolume , no,\n\n1962,\n\n33. Quine.) W# Wo0rd and\n\nMIT Press Cambridge, Mass.. 1960.\n\n34. Reichenbach, H. Elements of N.,Y.,i . 94 7\n\nic Lo\n\nThe Macmillan Co.,,\n\n35# Research Lab. of Electronics and Computation Center, MIT, COMIT ProgKammerls Reference Manual. MIT Press. Cambridge, Mass. 1961.\n\n36, Samuel, A. L \"Some Studies i, Machine Learning Using the Game of Gheckers)\" IBMj* of Research ad Development, lume 3 no. 3. 1959\n\n37c Shaw C* Jo \"JOVIAL and It!3 Documen:tation,\" Coram. ACM, Volume 3, no 6 1963.\n\n38* Simmons, R. F., et. al. \"Toward the Synthesis of Human Language\n\nB-iohavior,\" SPv\n\nSystems Development Corp., Santa Monica, Cal.\n\n39. Simon, H. A. \"Experiments with a heuristic oiler RAND Corp.,, Santa Monica, California, 1961.\n\nPaper P*2349,\n\n40. Slagle, J. \"A Computer Program for Solving Problems in Freshman Calculus,\" J. ACM. Jan., 1964 -and Doctoral Dissertation, Mathematics Department, MIT. May 1961.\n\n41. Solomonoff, R. J. \"An Inductive Inference Machine,\" IRE National Convention.,Record, pt 2 pp. 56-62* 1957.\n\n42. Sommersy F., T\n\nSemantic Structures and Automatic Clarification\n\nof Linguistic Ambiguity,\" International Electric Corp.) Paramus,\n\nRwJ. 0 1961.\n\n43. Supes, P. Introduction t;o L N.J. 1957.\n\nVan Nostrand Co., Princeton,\n\n44. Ullman, S. Words and Their Use.. Philosophical Library, NAY. 1951.0\n\n45. Walpole, H. R. Semantics-, The Nhature of Words and Their Meaningso W.W. Norton and Co. N.Y. 1941.\n\n46. Wang, H. \"Toward Mechanical Mathematics,` Development, Volume 4 no. 1. 1960.\n\nResearch and\n\n47,\n\niff,\n\nP6 Semantic- Anallys.is* Cornell U. Press, Ithaca,\n\nN.Y. 1960."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "student",
      "Paper": "Natural Language Input for a Computer Problem Solving System",
      "AtlasYear": 1964,
      "Status": "indexed",
      "Method": "windows-ocr-with-reviewed-citation-metadata",
      "SourceUrl": "https://bitsavers.trailing-edge.com/pdf/mit/ai/aim/AITR-219.pdf",
      "PdfSha256": "441D44A16C3624DEB317831E13303C7975F9288598B201125639F9FCFF27B978",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 1,
          "EndLine": 47,
          "PdfPages": [
            123,
            124,
            125,
            126
          ],
          "Text": "1. Berkeley, E.C. and D.G. Bobrow (eds.). The Programming Language LISP: Its Operation and Applications. Information International, 1964.\n2. Black, F. A Deductive Question-Answering System. Ph.D. thesis, Harvard University, 1964.\n3. Bobrow, D.G. METEOR: A LISP Interpreter for String Transformations. In reference 1.\n4. Bobrow, D.G. Syntactic Analysis of English by Computer: A Survey. Proceedings FJCC, 1963.\n5. Bobrow, D.G. and B. Raphael. A Comparison of List-Processing Computer Languages. Communications of the ACM, April 1964.\n6. Bobrow, D.G. and J. Weizenbaum. List Processing and the Extension of Language Facility by Embedding. Transactions IEEE PGEC, August 1964.\n7. Chomsky, A.N. On the Notion Rule of Grammar. Proceedings of the Symposium in Applied Mathematics, volume 12.\n8. Chomsky, A.N. Syntactic Structures. Mouton, 1957.\n9. Coffman, E.G., J.I. Schwartz and C. Weissman. A General-Purpose Time-Sharing System. Proceedings SJCC, April 1964.\n10. Cohen, D. Picture Processing in a Picture Language Machine. NBS Report 7885, April 1963.\n11. Coleman, M. A Program to Solve High School Algebra Story Problems. MIT term paper for 6.539, 1964.\n12. Cooper, W.S. Fact Retrieval and Deductive Question Answering. JACM 11(2), April 1964.\n13. Corbato, F.J. et al. The Compatible Time-Sharing System. MIT Press, 1963.\n14. Darlington, J. Translating Ordinary Language into Symbolic Logic. Memo MAC-M-149, MIT, March 1964.\n15. Feigenbaum, E. The Simulation of Verbal Learning Behaviour. In reference 16.\n16. Feigenbaum, E. and J. Feldman (eds.). Computers and Thought. McGraw Hill, 1963.\n17. Feldman, J. Simulation of Behaviour in the Binary Choice Experiment. In reference 16.\n18. Garfinkle, S. Heuristic Solution of First Year Algebra Problems. Working Paper 11, Management Science Group, University of California, Berkeley, 1962.\n19. Green, B.F., A.K. Wolf, C. Chomsky and K. Laughery. Baseball: An Automatic Question Answerer. Proceedings WJCC, May 1961.\n20. Harris, Z. Discourse Analysis. Language 28(1), January-March 1952.\n21. Harris, Z. String Analysis of Sentence Structure. Mouton, 1962.\n22. Kirsch, R.A. and B.K. Rankin III. Modified Simple Phrase Structure Grammar for Grammatical Induction. NBS Report 7890, 1963.\n23. Klein, S. and R.F. Simmons. Syntactic Dependence and the Computer Generation of Coherent Discourse. Mechanical Translation, 1963.\n24. Kuck, D. A Problem Solving System with Natural Language Input. Ph.D. thesis, Northwestern University, 1963.\n25. Kuno, S. and A. Oettinger. Syntactic Structure and Ambiguity of English. Proceedings FJCC, November 1963.\n26. Lamb, S.M. Outline of Stratificational Grammar. University of California, Berkeley, 1962.\n27. Lehman, W.P. and E.D. Pendergraft. Machine Language Translation Study No. 16. Linguistics Research Center, University of Texas, June 1963.\n28. Lindsay, R.K. Inferential Memory as the Basis of Machines which Understand Natural Language. In reference 16.\n29. Mathews, G.H. Analysis by Synthesis of Sentences in a Natural Language. First International Conference on Machine Translation and Applied Language Analysis, HMSO, 1962.\n30. McCarthy, J. Programs With Common Sense. Symposium on Mechanization of Thought Processes, HMSO, 1959.\n31. McCarthy, J. et al. LISP 1.5 Programmers Manual. MIT Press, 1963.\n32. Minsky, M. Steps Toward Artificial Intelligence. In reference 16.\n33. Morris, C.W. Foundations of the Theory of Signs. International Encyclopedia of Unified Science 1(2), University of Chicago Press, 1955.\n34. Newell, A. et al. Report on a General Problem Solving System. International Conference on Information Processing, UNESCO House, Paris, 1959.\n35. Ogden, C.K. A System of Basic English. Harcourt-Brace, 1934.\n36. Phillips, A.V. A Question-Answering Routine. Master's thesis, MIT Mathematics Department, 1960.\n37. Quine, W.V. Word and Object. MIT Press, 1960.\n38. Raphael, B. SIR: A Computer Program for Semantic Information Retrieval. Ph.D. thesis, MIT Mathematics Department, 1964.\n39. Sillars, W. An Algorithm for Representing English Sentences in a Formal Language. NBS Report 7884, April 1963.\n40. Simmons, R.F. Answering English Questions by Computer: A Survey. SDC Report SP-1556, April 1964.\n41. Simmons, R.F., S. Klein and K.L. McConlogue. Indexing and Dependence Logic for Answering English Questions. American Documentation, in press.\n42. Skinner, B.F. Verbal Behaviour. Appleton Century Croft, 1957.\n43. Walker, D.E. and J.M. Bartlett. The Structure of Language for Man and Computers: Problems in Formalization. First Congress on the Information Sciences, Vista Press, 1963.\n44. Yngve, V. A Model and an Hypothesis for Language Structure. Proceedings of the American Philosophical Society 104(5), 1960.\n45. Yngve, V. COMIT Programmers Reference Manual. MIT Press, 1961.\n46. Yngve, V. Random Generation of English Sentences. Proceedings 1961 International Conference on Machine Translation and Applied Language Analysis 1, HMSO, 1962.\n"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Scanned bibliography pages processed with Windows OCR. The readable citation metadata was checked against the images and OCR; raw OCR is retained with page numbers and hashes. Spelling and source-publication inconsistencies may remain. Only reviewed Atlas matches become citation links."
      ],
      "OcrPages": [
        {
          "PdfPage": 123,
          "Text": "(1)\n(2)\n(3)\n(12)\nBIBLIOGRAPHY\nBerkeley, ETC. and Bobrow (eds.3, Erczramninz\nInfortne—\nIts Operation and Aoplleaticng,\n1964.\ntion International, , C8mbridgei Ya$$.;\nDeductive Question—kvi$vering System,\" Ph. D.\nBlock* P. ,\nNest* i Division of Engineering and Applied Physics, Harvard\n1964.\nUniversLty, Cambridge , Maas.\nD.c., iiXETEOR: A LISE Interpreter for string\nfornatioaa,ii In\nBobrov, fisyneactLc Analysis of English by\nA Survey, ii Proc. FJCC, Spartan Press , Baltimore. Md; 1963\nBobrows D.G. end B. Raphsel„ Compari$on of List—Pfoeegsitig\nCornpuEer A-(947 l,\nD.C. and J. \"List proeegging and the\nExtension Language Foeilfty by Embedding,\" Trans. IEEEs\nPGECi August, 1964.\nChomsky, , the Notion 'Rule of Granma\"'\nof the in AppJ1ed vol. 12.\nChomsky, Syntactic Structures, Mout:aa and Co.\nhege; 1957.\nCoffman, E.G., Schwerte End C. geissman, \"A General—Por—\npose System.\" Proc. SJCC, SpsEtan Press,\ntiture. Yd.; April, 1964.\nCohen, \"Picture Troce\"fng in a Picture Language Machine\n1963.\nNBS Report 7885, nepE. ot Coumerce. Wash. , D.C.; Apr LI,\nColeman, M. , Program eo Solve High school Algebra Story\nProblems , \" Tern Paper for 6.539; 1964.\nCooper, K. S. * \"Eact Retrieval end neduceive Question AnswerLug,\nvol. no. 2; April, 1964.\ntible Time-Sheri\ncorbato, r.J., et\n1963.\nMIT Press, Cambridge,",
          "Receipt": {
            "TextSha256": "0700D77CDF9202BF65AFDB4A7591423FA8133430CAEE8D44E60F8D1C3FE43E89",
            "ProcessedAtUtc": "2026-09-16T22:50:34.3714036Z",
            "ImageSha256": "E968906AD6C2C6CE0850B22E3E078CEBB24FECB2B33FC1F2148823B049BD93A8",
            "Language": "en-US",
            "Lines": 42,
            "Engine": "Windows.Media.Ocr",
            "Output": "student-refs-000123.ocr.txt",
            "Image": "student-refs-000123.png"
          }
        },
        {
          "PdfPage": 124,
          "Text": "( 18)\n(20)\n(21)\n(22)\n(26)\n(25)\n(26)\n(27)\nDarlington, J x, Ordinary Language Sym.\nbolie Logic/' Memo W,C-M-149, Project %4C, MIT; March, 1964.\nFeigenbaum, E. S of Verbal\niaur,\" in (16).\nFeigenbaum, E. and . FeldiL2n (edE . ) and Ihouzht\nHill, *grk; 1963.\nFeldnan. J. , \"Simulation Of Behaviour in the Choice\nExperiment , \"\nGarfinkle,\nsolution of First Year Algebra\nProblernG , •i Working paper t l, Manegement\nGroup i University of California, Berkeley, Cali L;\nGreen, B. F, , A.K. Wolf* C. Chomsky and K. Laughery,\n\"baseball: An Question Answerer, *JCC;\nmay, 1961.\nHarris. Z.\n\"Discourse Lanskuaee, vol. 28, gg.\n- Match, 1952.\nHarris, Z. * String *rue lysis of Sentence Structure,\nend Co., The 1962.\nKirsch, R.A. and B.K. Rankin \"Yodified Simple Phrase\nSC Gram.rnars far Grermatieal Induction,\" NBS Report\nDept. Conmetce, Wesh., D ; 1963.\nKleins S. •tbd Sirrunons, iiSyntactiC Dependence end the\nConputer Generation of Cob.ergac Discourse , if Mechanical\nKuck,. D. , \"A Problem Solving Syseet *ith Natur4L Language\nPh. D. Thesis, Technological Institute,\nUniv.„ IL linoLs; 1963.\nXur:g,. and Oettinger, • 'Syntactic Structure and\nbiggity English,\" Proc. P.JCC, Spartan Press, Ba\n1-\ntiäare, W. ; Ko-v., 1963.\nS.H+* Outline of $ceacificational Gremar, univer-\nOf California, Berkeley,\n1962 .\nLehman, W.P. and E.D. Pendergrsft, Heehine Lonzuaae Trans-\nStudy Linguistic Reseerch Center, Univer—\nGity of TeK.eE. Texas; June. 1963.",
          "Receipt": {
            "TextSha256": "D0EDA5F2F88161FFC74CBA36532909B1AD8289E3AE6295E16F1C2F50D0CDAC8B",
            "ProcessedAtUtc": "2026-09-16T22:50:34.5256863Z",
            "ImageSha256": "F476386166F68FC7A9E7656E52672DBAEA2ACA5953BFD7992FCCF832CEE92086",
            "Language": "en-US",
            "Lines": 46,
            "Engine": "Windows.Media.Ocr",
            "Output": "student-refs-000124.ocr.txt",
            "Image": "student-refs-000124.png"
          }
        },
        {
          "PdfPage": 125,
          "Text": "(28)\n(29)\n(30)\n01)\n(32)\n(33)\n( 34)\n(35)\n(16)\n(37)\n(38)\n(39)\nLindsey, RX. , \"Inferential Mengry as Che\nWhich Understand Nature 1 Language, ii in\nMathews, \"An,eLyEiE by Synthesis of Sentences La\nNatural Language,\" First International Conference on\nhöchine Trans Letion end\nlied Lan e Anal s Ls, KMSO,\nLondon; L9S2.\nMcCarthy, J. , i Tcograxns With Cotunon Sense,\" Proc. of the\non Yechanization of thoueht Processes,\nLondon ; 195 9'.\nMcCarthy. J. i et al., LISP 1.5 Pro Manuel, MIT\nPress, Yass. ;\nMinsky. M.. \"Steps Tovard Artificial Intel ligence,l' in (16) .\nMetris, C.W. ,\n• iFoundations the Theory of Signs,\" Inter—\nnat ion.\"l Encvclooediö of Unified vol. 1, no. 2,\nUniversity of Chic.ego Press T Chicago;\nNevells A. et \"Report oa CeneraL problem Solving\nSystem,\" Proc.\nInternationel Conference Jnfgrmetlgn\ntnusco 1959.\nOgden, C.K., System of Basic Enhlish, Harcourt—brace,\nNeu York;\nPhillips, Routine i \" Thesis,\nMathematics Department, Cambr Mass . ;\nQuine, W.V., Word and Obleec, Press, Cfihbridge,\n1960 .\nRaphael, B,. , \"'SIR: Coeputer Frograrn for Sernnntie\ntion Retrieval , • i\nPh. D. I%æsis, Nethernatic8 Department\nCambridge. Yass . ; 1964\nSillars,\n• •An Algorithm for Representing English\nin a Formal Lenguege,•• BBS Report 1884, Department of Cota—\nmerce, Wash. , D. C.i April, 1963.\nSinnons, , \"Answering English Questions by Computer—\nA Survey,\" SDC Report SP- 1556, Sance Monica, CA lif.; April,\nSignor's , S. Klein ond K.L. McCon010g•Lm \"Indexing\nand Dependency i\nAmerican Doeu:nentation, (in press)\n126\n1964.",
          "Receipt": {
            "TextSha256": "2739A880D26E7190EA6FFB9017EC0B2BABBD7A1FBFAB88F48388EB71B82DB154",
            "ProcessedAtUtc": "2026-09-16T22:50:34.6856633Z",
            "ImageSha256": "798093864688AAE79647584BB3F8A10060D6753495A6E480769DA414CAB8F5E3",
            "Language": "en-US",
            "Lines": 54,
            "Engine": "Windows.Media.Ocr",
            "Output": "student-refs-000125.ocr.txt",
            "Image": "student-refs-000125.png"
          }
        },
        {
          "PdfPage": 126,
          "Text": "(4-2)\n( 43)\n(45)\n(46)\nSkinner, Verbal Behaviour, Appleton Century Croft,\nYork; 1957.\nWalker, D.E. and Bart leet, Structure of Languege\nfor Man FrobLÄS in Formalization, Proc .\nF on the Information Sciences, Viste Press;\n1963.\nYagve, Model and an Hypothesis for Language\nStructure S Of the American philoso hfceL\n, Prograttners Reference Manual,\nMir ,\nYngve, V.. Generatioa of English SentenceS\nPrac. 1961 Conference On Mochine Trens•\nl•tLon and Language Ana lysis, vol. l, HMSO\nLondon; 1962,\n127",
          "Receipt": {
            "TextSha256": "AA93437F18CFFD590066A7808190D20016FE6BF71D669CBBFCEE682BCC3EFDEE",
            "ProcessedAtUtc": "2026-09-16T22:50:34.7846692Z",
            "ImageSha256": "25FEA862E7CD0BBD4D28101280920516BAA87CC646C6AED4AE2D0B76A6FC7EC7",
            "Language": "en-US",
            "Lines": 19,
            "Engine": "Windows.Media.Ocr",
            "Output": "student-refs-000126.ocr.txt",
            "Image": "student-refs-000126.png"
          }
        }
      ]
    },
    {
      "Slug": "resolution",
      "Paper": "A Machine-Oriented Logic Based on the Resolution Principle",
      "AtlasYear": 1965,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://web.stanford.edu/class/linguist289/robinson65.pdf",
      "PdfSha256": "757C6BCEF2E07FDAE1E604CEA0556413B82AD46F3F26D62DBC20F30675B9886B",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 401,
          "EndLine": 405,
          "PdfPages": [
            19
          ],
          "Text": "1. C~ultcu, A. A note oa the Entscheidungsproblem. J. S~tmb. Logic 1 (1936), 40-41. Correction, ibid., 101-102.\n;. Daws, M., aaD Pb'q'N~, H. A computing procedure for quantification theory. J. AC.,lf 7 (Mar. 1960), 201--215.\n3. FetEr)X*AN,J. A semi-decision procedure for the functi(mal calculus. J. ACM 10 (Jan. 1963), 1-24.\n4. G~,~roa~, P. C. A proof method for quantifieatioa theory. I B M J. [lea. Develop. 4 (1960), 28-35.\n5. ROBINSON,J. A. Theorem-proviag on the computer. J. ACM 10 (Apr. 1963), 163-174."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "eliza",
      "Paper": "ELIZA: A computer program for the study of natural language communication between man and machine",
      "AtlasYear": 1966,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://web.stanford.edu/class/cs124/p36-weizenabaum.pdf",
      "PdfSha256": "D0C58989BFF11741AEC3FBF52A49B3DCD0EDE45EBABFA249815602B81EF07C1A",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 758,
          "EndLine": 763,
          "PdfPages": [
            8
          ],
          "Text": "1. ABEl.SON, Pt. P., AND CARROLL, J. D. Computer simulation of individual belief systems. Amer. Behav. Set. 9 (May 1965), 24-30.\n2. Goax, S. Semiotic relationships in ambiguously stratified bmguage systems. Paper presented at Int. Colloq. Algebraic Linguistics and Automatic Theory, Hebrew U. of Jerusalem, Aug. 1964.\n3. Bom¢ow, D . G . Natural language input for a computer problem solving system. Doctoral thesis, Math. Dept., MIT. Cambridge, Mass., 1964.\n4. WEIZENB_~UM,J. Symmetric list processor. Comm. A C M 6, (Sept. 1963), 524-544.\n5. ROGEaS, C. Client Centered Therapy: Current Practice, Implications and Theory. Houghton Mifflin, Boston, 1951.\n6. YNGVE, J. COMIT Programndng Manual. M[T Press, Cambridge, Mass., 1961."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "semantic-memory",
      "Paper": "Semantic Memory",
      "AtlasYear": 1966,
      "Status": "indexed",
      "Method": "pdf-text-and-windows-ocr",
      "SourceUrl": "https://archive.org/download/DTIC_AD0641671/DTIC_AD0641671.pdf",
      "PdfSha256": "C3EA7BF12C2A1B6F33D0087E71595F76A44C761DF312D6136721AFD1B8912158",
      "Sections": [
        {
          "Section": "Bibliography, printed pages 166-175",
          "PdfPages": [
            175,
            176,
            177,
            178,
            179,
            180,
            181,
            182,
            183,
            184
          ],
          "Text": "I\r\n\r\nI\r\n\r\nI BIBLIOGRAPHY\r\n\r\nBanerji, R. B. A language fc the description of concepts. Un-\r\n\r\nI\r\n\r\nI published dittoed paper. Systems Research Center, Case Insti-\r\ntute uf Technology, 1964.\r\n\r\nBartlett, P. C. Remembering, a study in experimental and\r\n\r\nI\r\n\r\nsocial psychology. Cambridge, England: Cambridge University\r\n\r\nPress, 1932. I\r\n\r\nBaylor, G. W., and Simon, H. A. A chess mating combinations program. Proceedings of Spring Joint Computer Conference,\r\n\r\n1966, 28, 431-447.\r\n\r\nBerkeley, E. C, and Bobrow, D. G. The programming language LISP: its operation and applications. Cambridge, Massa chusetts: Information International, Inc., 1964.\r\nBobrow, D. G. Syntactic analysis of language by computer — a survey. Proceedings of the Pall Joint Computer Conference, 1963,. 24, 365-387.\r\n\r\nBobrow, D. G. Natural language input for a computer problem\r\n\r\nsolving system. Unpublished PhD. dissertation, M.I.T. 1964.\r\n\r\nr\r\n\r\nAlso Project MAC, Report #TR-1, 1964, Cambridge, Mass.\r\n\r\nI\r\n\r\nBobrow, D. G. and Tc-itelman, W. Format-directed List Processing in LISP. Cambridge, Massachusetts: Bolt Beranek and Newman Report #1366, 1966.\r\n\r\n!\r\n166\r\n1\r\n\r\n.■gjyfJ.-fgiag'-'.^-a»- - •-^■■. ^-1»-....-»—■ ■^—\r\n\r\n, -\r\n\r\n■■ ■ ■ \"■-' ■■ —'\r\n\r\n•my<m •*- ■\r\n\r\n■' ■ HII-^W ji'jji IMWIII in j | p^MI -■UgU *\r\n\r\n\nBruneis J. S. On perceptual readiness. Psychological Review, 1957. 64, 123-152.\r\nBruner, J. S., Goodnow, J. J., and Austin, C. A. A study of thinking. New York: John Wiley and Sons, Inc., 1956.\r\nBruner, J. S. and Minturn, A. L. Perceptual identification and perceptual organization. Journal of Genetic Psychology, 1955. 53, 18-21.\r\nChomsky- N. Aspects of the theory of syntax. Cambridge, Massachusetts: The M.I.T. Press, 1965.\r\nChomsky, N. and Miller, G. A. Introduction to the formal analysis of natural languages. In D. R. Luce, R. R. Bush, and E. Galanter (Eds.), Handbook of Mathematical Psychology, Vol. II. New York: John Wiley and Sons, 1963.\r\nChcmsky, N. Review of Skinner, B. F. Verbal Behavior. Language, 1959. 35, 26-58.\r\nCliff, N. Adverbs as multipliers. Psychological Review, 1959, 66, 27-44.\r\nCreeiman, M. B. The experimental investigation of meaning. New York: Springer Publishing Company, 1966.\r\nDarlington, J. Translating ordinary language into symbolic logic. Memorandum MAC-M-149, Project MAC, M.I.T., Cambridge, Massachusetts, 1964.\r\nDeese, J. On the structure of associative meaning. Psychological Review, 1962, 69 (3), 161-175-\r\n\r\nn\r\n\r\n167\r\n\r\n^JLJH\r\n\r\n■^P1^^!^^'\r\n\r\n—i M ^ ■ > ■\r\n\r\n\"S^P\r\n\r\n-^PB\r\n\r\n\nErvin, S. M. Changes with age in the verbal determinants of\r\n\r\nword association. American Journal of Psychology, 1961,\r\n\r\n|\r\n\r\n74, 361-372.\r\n\r\nFeigenbaum, E. A., and Simon, H. A. Performance of a reading\r\n\r\nI\r\n\r\nI task: by an elementary perceiving and memorizing program.\r\nBehavioral Science, 8 (l), 72-76, 1963\r\n\r\nFeigenbaum, E. A. An information processing theory of verbal\r\nI learning. Report P-l8l7^ the RAND Corporation, Santa Monica,\r\nCalifornia, 1959-\r\nFeigenbaum, E. A. and Peldman, J. Computers and Thought, New York: McGraw Hill, 1963.\r\n\r\nFifth Annual Report, the Center for Cognitive Studies, 1964-65» The Center for Cognitive Studies, Harvard University, Cambridge, Massachusetts, 1965.\r\n\r\nPillmore, C. J. A proposal concerning English prepositions. Paper presented at M.I.T., Cambridge, Massachusetts, April, 1966.\r\n\r\nFunk and Wagnall's new \"standard\" dictionary of the English language. New York: Funk and Wagnalls Company, 1959.\r\n\r\nGelernter, H. Hansen, J.R. and Loveland, D.W. Empirical ex-\r\nplorations of the geometry-theorem proving machine.\r\nI Proceedings of the Western Joint Computer Conference, i960,\r\n17, 143-147.\r\n1\r\n\r\nI\r\n\r\nI\r\n168\r\nI\r\n\r\n1 w J\"^^!\"-'-=«PrTr -V^'-S Jt\r\n\r\n*\r\n\r\n;\r\n\r\n'—1 ■■»«-■- -= ■ \"*■—\r\n\r\n- ■ ■_ ——\r\n\r\n;\r\n\r\n'---■ '^P» 1 P--i — ■■-■ jf—jjM^ynüiiiMpiiMjM^^By^p,\r\n\r\n\nGreen, E. P., Wolf, A. K., Chomsky, C, and Langhery, K. Baseball: An automatic question answer. Proceedings '>£ the Western Joint Computer Conference, 1961, 19^ 219-224.\r\nHalle, M., and Stevens, K. N. Speech recognition: A model and a program for research. In Podor, J. A., and Katz, J. J. (Eds), The Structure of Language; Readings in the Philosophy of Language. Englewood Cliffs, New Jersey: Prentice-Hall, Inc., 1964, 604-612.\r\nHays, D. G., (Ed). Readings in automatic language processing. New York: American Elsevier Publishing Co., 1966.\r\nHunt, E. B, Concept learning: an information processing problem. New York: John Wiley and Sons, Inc., 1962.\r\nKatz, J. J. and Poder, J. A. The structure of a semantic theory. Language, 1963, 39, 170-210.\r\nKatz, J. J. and Postal, P. M. An Integrated theory of linguistic descriptions. Cambridge, Mass. The M.I.T. Press, 1964.\r\nKelly, G. The psychology of personal constructs; Volume I. New York: W. W. Norton and Company, Inc., 1955.\r\nKlein, S. Automatic paraphrasing in essay format. . SP-l602/00l/00, System Development Corporation, Santa Monica, California, 1964.\r\nKlein, S. and Simmons, R. P. Syntactic dependence and the computer generation of coherent discourse. Mechanical Translation, 1963, 7(2), 50-61.\r\n169\r\nfmrn 'i-i\r\n\r\n\nKuno, S. K. Multiple-Patn Syntactic Analyzer. Mathematical linguistics and automatic translation. Report No. N3F-8, 1963, Computation Laboratory of Harvard University, Cambridge, Massachusetts.\r\n\r\nKuno, S. K. The predictive analyzer. Communications of the Association for Computing Machinery, 1965, 8 (7). 453-^62. Reprinted in Hays, 1966.\r\n\r\nLakoff, G. On the nature of syntactic irregularity. Ma the matical Linguistics and Automatic Translation. Report No. NSP-16, 1965, Computation Laboratory of Harvard University, Cambridge, Massachusetts.\r\n\r\nLamb, S. The sememic approach to structural semantics. In K. A. Romney and R. D'Andrede (Eds), Transcultural studies in cognition. American Anthropologist, 1964, 66 {3), Part 2.\r\n\r\nLane, H., and Schneider, B. Some discriminative properties of syntactic structures. Journal of Verbal Learning and Verbal Behavior, 1963, 2(5-6), 457-461.\r\n\r\nLindsay, R. K. Inferential memory as the basis of machines\r\n\r\n1\r\n\r\nwhich understand natural language. In Feigenbaum, E., and\r\n\r\nPeldman, J. (Eds.) Computers and Thought. New York:\r\n\r\nMcGraw Hill Books Co., 1963> 217-233.\r\n\r\nMatthews, G. H. Analysis by synthesis of sentences of natural languages. Proceedings of 1st Int. Gong, on machine translation of languages and applied language analysis, 1961, Teddlngton, England: National Physical Laboratory,\r\n\r\n170\r\n\r\nI\r\n\r\nf\r\n\r\n\nMcCarthy, J., Abrahams, P, W., Edwards, D. J., Hart, T. P. and Levin, M. I. LISP 1.3 Programmer's Manual, Cambridge, Massachusetts: the M.I.T. Computation Center, 1962.\r\nMiller, G, A. The magical number seven, plus or minus two: some limits on our capacity for processing information. Psychological Review, 1956, 63.»81-96.\r\nMiller, G. A., and Chomsky, N. Finite models of language users. In Luce, R. D., Bush, R, L,, and Galanter, E., (Eds). Handbook of mathematical psychology. Volume II, New York: John Wiley and Sons, Inc., 1963, 419-488.\r\nMiller, G. A., Galanter, E., and Pribram, K. H. Plans and the structure of behavior. New York: Holt, I960.\r\nMinsky, M. Steps toward artificial intelligence. Proceedings of the Institute of radio engineers, 1961, 49j 8-30.\r\nMinsky, M. A selected descriptor-indexed bibliography to the literature on artificial intelligence. In Feigenbaum, E. A., and Feldman, J. (Eds.) Computers and Thought. New York: McGraw-Hill Book Co., Inc., 1963, 453-456.\r\nNewell, A. (Ed.) IPL-V programmer's reference manual. Memorandum RM-3739-RC, The RAND Corporation, Santa Monica, California, 1963.\r\nNewell, A., Shaw, J. C. and Simon, H. A. The processes of creative thinking. In H. E. Gruber, G. Terrellj and M. Wertheimer (Eds.) Contemporary approaches to creative thinking, New York: Atherton Press, 1962, 63-119-\r\n\r\nTT\r\n\r\n171\r\n\r\n^BSHWP\r\n\r\n-1 i ■- — J- mmm\r\n\r\n\nI\r\n\r\nI Newell, A., Shaw, J. C, and Simon, H. A. Chess playing proI grams and the problem of complexity. I.B.M. Journal of\r\nResearch and Development, 1958, 2, 4, 320-335.\r\n\r\nNewell, A., and Simon, H. A. Computers in psychology. In\r\n\r\nI\r\n\r\nR. D. Luce, R. Bush, and E. Galanter (Eds.), Handbook of\r\n\r\nMathematical Psychology, Volume 1. New York: John l-illey\r\n\r\nand Sons, 1963, 361-408.\r\n\r\nNewell, A., and Simon, H. A. An example of human chess play in the light of chess playing programs. Pittsburgh, Pennsylvania: Carnegie Institute of Technology, 1964, (dittoed).\r\n\r\nOgden, O.K. The General Basic English Dictionary. New York: W. W. Norton and Company, Inc., 1942.\r\n\r\na\r\n\r\nOlney, J. Building a concept network for retrieving informa-\r\n\r\ntion from large libraries: Part I. Technical Memorandum\r\n\r\nTM-634/OOI/II, System Development Corporation, Santa Monica,\r\n\r\nCalifornia, 1962.\r\n\r\n*\r\n\r\nOlney, J. C. Some patterns observed in the contextual specialization of word senses. Information Storage and Retrieval, 1964, 2, 79-101.\r\n\r\nOsgood, E. C, Suci, G. J. and Tannenbaum, P. H. The measurement of meaning. Urbana, Illinois: University of Illinois Press, 1957.\r\n\r\nOsgood, C. E. On understanding and creating sentences. American Psychologist, 1965, 18, 735-751.\r\n\r\n172\r\n\r\n1\r\n\r\nI\r\n\r\n!-\"^^y!:_J.i--#-j-»ij»i J- ■' 1. i_\r\n\r\n■-- \" -\"■-\r\n\r\n-»=—\r\n\r\n?——\r\n\r\nJll I\r\n\r\n1 liumpi—^■111 ■■THHmn\r\n\r\n\nPaige, J. J., and Simon, H. A. Cognitive processes in solving algebra word problems. In Kleinmuntz, B. (Ed.), Problem Solving: Research, Method and Theory. New York: John Wiley and Sons, 1966.\r\nPetrick:, S. R. A recognition procedure for transformational grammars. Unpublished Ph.D. dissertation, M.I.T., Cambridge, Massachusetts, 1965.\r\nPlaget, J. The psychology of intelligence. Translated by M. Cook and D. E. Berlyne. London, England: Routledge and Kegan Paul, 1950.\r\nQuillian, R. A design for an understanding machine. Paper presented at a colloquium: Semantic problems, in natural language. King's College, Cambridge University, England, September, 1961.\r\nQuillian, R. A revised design for an understanding machine. Mechanical Translation, 1962, 7, 17-29. (a)\r\nQuillian, R. A semantic coding technique for mechanical English paraphrasing. Internal Memorandum of the Mechanical Translation Group, Research Laboratory of Electronics, M.I.T., Cambridge, Massachusetts, August, 1962. (b)\r\nQuillian, R. A notation for representing conceptual information: an application to semantics and mechanical English paraphrasing. SP-1395, System Development Corporation, Santa Monica, California, 1963.\r\n\r\nW\r\n\r\n173\r\n\r\n^P^^IHJJ_»_IL-J»_- iHI-L „ J L ■ -4^-XJI.\r\n\r\nmi n J^__ M _\r\n\r\n\nQuillian, R., Wortman, P. and Baylor, G. W. The programmable Plaget: behavior from the standpoint of a radical computerist. Unpublished dittoed paper, Carnegie Institute of Technology, 1965.\r\n\r\nRaphael, B. A Computer program which \"understands.\" Proceedings of AFIPS, 1964, Pall Joint Computer Conference, PP. 577-589.\r\n\r\nRazran, G. H. S. A quantitative study of meaning by a. conditioned salivary technique (semantic conditioning). Science, 1939a, 90, 89-90 [102, 106].\r\n\r\nReich, P. A. A stratlflcatlonal theory of language acquisition. Working Paper #4 (IP-4), Department of Psychology and Mental Health Research Institute, University of Michigan, Ann Arbor, Michigan, 1966.\r\n\r\nReld, L. S., Henneraan, R, H. and Long, E. R. An experimental analysis of set: the effect of categorical restriction. American Journal of Psychology, i960, 73, 568-572.\r\n\r\nRelss, R. P. An abstract machine based on classical association psychology. Technical Memorandum of Llbrascope Division, General Precision, Inc., Glendale, California, 1961.\r\nI Reitman, W. R. Cognition and thought; an information pro-\r\ncessing approach. New York: John Wiley and Sons, Inc., 1965.\r\n\r\nI\r\n\r\nI\r\n\r\n174\r\n\r\nI\r\n\r\nI\r\n\r\n■ «■■■■■■Himg! ..■jM.Ji..,.,. ,\r\n\r\n_-_-\r\n\r\n,\r\n\r\n.-___-_\r\n\r\n. - ^—g;...\r\n\r\nf ^.p.\r\n\r\n\nRubensteln, H. Problems in automatic word disambiguation. Paper presented at a conference on Computer-Aided Semantic Research^ held at Las Vegas., Nevada, December, 1965.\r\nSimmons, R. P. Synthetic language behavior. Data Processing Management, 1963, 5 (12), 11-18.\r\nSimon, H. A. and Feigenbaum, E. A. An informationprocessing theory of some effects of similarity, familiarization, and meaningfulness in verbal learning. Journal of Verbal Learning and Verbal Behavior, 1964, 3, 385-397.\r\nSimon, H. A. and Kotovsky, J. Human acquisition of concepts for sequential patterns. Psychological Review, 1963.» 70, 534-546.\r\nSkinner, B. F. Verbal Behavior. Appleton Century Crofts, Inc., New York, 1965«\r\nYngve, V. H. A model and an hypothesis for language structure. Proceedings of the American Philosophical Society, i960, 104, 444-466.\r\nYngve, V. H. Comit Programmer's Reference Manual. Cambridge, Massachusetts: M.I.T, Press, 1961.\r\n\r\n175\r\n\r\n■p—■—1\r\n\r\nILL-'U ■\r\n\r\n^JJIUB« 11\r\n\r\n.jML.^^ijuagi"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "The complete ten-page bibliography is indexed from Banerji through Yngve. The source scan was recovered from DTIC AD0641671. Windows OCR was run on all bibliography pages and checked against the existing text layer; both are preserved. The following appendix is excluded. OCR spelling and layout errors may remain."
      ],
      "BibliographyEdition": "1966 dissertation/report, DTIC AD0641671",
      "OcrPages": [
        {
          "PdfPage": 175,
          "Text": "BIBLIOGRAPHY\nBanerji, R. B. A language fc the description of concepts. Ün-\npublished dittoed paper, Systems Research Centel , Case Insti—\ntute of Technology, 1964.\nRemembering, a study In experimental and\nBartlett, F. C.\n1\nsocial psychology. Cambridge, England:\nPress, 1932.\nCambridge University\nA chess mating combinations\nBaylor, G. W. , and Simon, H. A.\nProceedings of Spring Joint\nprogram .\n1966, 28, 431\nBerkeley, E. C. , and Bobrow, D. G. The\nIts operation and applications.\nLISP :\nComputer Conference,\nprogramming language\nCambridge, Massa\n1964 .\nIne . ,\nchusetts:\nBobrow, D.\nsurvey .\n1963, 24,\nInforma tion Interna tional,\nG.\nSyntactic analysis of language by compu ter\na\nProceedings of the Fall Joint Computer Conference,\n365-387 •\nBobrow, D. G. Natural language Input for a computer problem\nsolving system. Unpublished PhD. dissertation, M.I.T. 1964\nAlso Project MAC, Report #QR-I, 1964, Cambridge, Mass.\nBoorow, D. G . and Teitelman, W. Format-directed List Process-\nIn LISP.\nCambyldge, Massachusetts: Bolt Beranek and\nNewman Report #1366,\n966 .\n166",
          "Receipt": {
            "TextSha256": "B247CEA8903A6C4BEF7E8F7ABDD85079F7FA3A7E15F2ACAA54BB593E96E5A3DB",
            "ProcessedAtUtc": "2026-09-16T23:48:49.5442549Z",
            "ImageSha256": "DCC098CECFB0DAB4BF14460323E80416B9313962F7F1032BCC485003CFDA828C",
            "Language": "en-US",
            "Lines": 42,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000175.ocr.txt",
            "Image": "semantic-memory-refs-000175.png"
          }
        },
        {
          "PdfPage": 176,
          "Text": "Bruner, J.\n1957, 64,\nBruner, J .\nthinking.\nBruner, J.\nS. On perceptual readiness.\nPsychological Review,\n123-152.\nGoodnow ,\nNew York:\nJ. J. , and Austin, C.\nJohn Wiley and Sons,\nInc .\nA study of\n1956 .\nS. cxnd Minturn, A . L. Perceptual\nidentification\nand perceptual organization .\nJournal of Genetic Psychology,\n1955, 53, 18-21.\nChomsky, N.\nAspects of the theory of syntax.\nThe M.I.T. Press, 1965.\nMassachusetts:\nCambridge ,\nChomsky, N. and Miller, G. A.\nIn troduc tlon to the formal\nanalysis of natural languages.\nIn D. R. Luce, R. R. Bush,\nand E. Galanter (Eds.), Handbook of Mathematical Psychology,\nVol. II. New York: John Wiley and Sons, 1963.\nChcmsky, N. Review of Skinner, B.\nLanguage, 1959, 35, 26-58.\nCliff, N . Adverbs as multipliers.\n66, 27-144.\nF. Verbal Behavior.\nPsy_chological Review, 1959,\nCreelman, M. B. The experimental investiga tion of meaning.\nNew York: Springer Publishing Company, 1966 .\nDarlington, j. Transla ting ordinary language into symbolic\nMemorandum MAC-M-149, Project MAC, M.I.T.,\nlogic .\nMassachuse tts, 1964.\nDeese, J. On the struc ture of associative meaning.\nlogical Review, 1962, 69 (3), 161-175.\n167\nCambridge ,\nPsycho-",
          "Receipt": {
            "TextSha256": "BFD0BF6E3293D9E8617FD0148C8EA80FFE6A5242FBFAAD023FF41DF20B9CF8C9",
            "ProcessedAtUtc": "2026-09-16T23:48:49.6940301Z",
            "ImageSha256": "D8F5DC5D26B915C566D7F0B61505CA2186EE0FAB06B1ABF22B251F86497593E6",
            "Language": "en-US",
            "Lines": 48,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000176.ocr.txt",
            "Image": "semantic-memory-refs-000176.png"
          }
        },
        {
          "PdfPage": 177,
          "Text": "Ervin, S. M. Changes with age In the verbal determinants of\nAmerican Journal of Psychology, 1961,\nword association.\n74, 361-372.\nperformance of a reading\nFeigenbaum, E. A. , and Simon, H. A.\ntask by an elemen tary perceiving and memorizing program .\n(1), 72-76, 1963\ninformation processing theory of verbal\nlearning. Report P -1817, the RAND Corporation, Santa Monica,\nBehavioral Science,\nFeigenbaum, E. A.\nAn\nCalifornia, 1959 •\nFeigenbaum, E. A. and\nFeldman, J.\nComputers and Thought,\nNew York:\nFifth Annual\nThe Center\nCambridge ,\nFillmore, C.\nMcGraw Hill, 1963.\nReport, the Center for Cognitive Studies, 1964-65.\nfor Cognitive Studies, Harvard University,\nMassachusetts, 1965 .\nJ. A proposal concerning English prepositions.\nPaper presented at M.I.T., Cambridge, Massachusetts,\nApril, 1966 .\nFunk and Wagnall 's new t' standard\" dictionary of the English\nlanguage. New York: Funk and Wagnalls Company, 1959 .\nGelernter, H. Hansen, J . R. and Loveland, D.W. Empirical ex—\nplorations of the geometry-theorem proving machine .\nProceedings of the Western Joint computer Conference, 1960,\n17, 143-147.\n168",
          "Receipt": {
            "TextSha256": "DD79E8B107D37CF82F84507388CC660459B4C0CF76312B06A735D178E1A71CDF",
            "ProcessedAtUtc": "2026-09-16T23:48:49.8301103Z",
            "ImageSha256": "73CEA80D123FB1A3603A3F9C1651B609F84B4C9AFEF2873CD9954AB3FD123E7D",
            "Language": "en-US",
            "Lines": 36,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000177.ocr.txt",
            "Image": "semantic-memory-refs-000177.png"
          }
        },
        {
          "PdfPage": 178,
          "Text": "Wolf, A. K. , Chomsky, C. ,\nand Langhery, K.\nGreen, B. F. ,\nAn automatic question answer. Proceedings\nBaseball :\nthe Western Joint Computer Conference, 1961, 19, 219-24.\nHalle, M. , and Stevens, K. N. Speech recognition: A model\nand a program for research. In Fodor, J. A . ,\nand Katz, J. J.\n(Eds), The Structure of Language:\nReadings in the Philosophy\nEnglewood Cliffs, New Jersey: Pren tice-Ha11,\nof Language.\n1964, 604 —612 .\nInc.,\n(Ed) .\nReadings in automa tic language processing.\nHays, D. G. ,\n1966 .\nAmerican Elsevier Publishing Co.,\nNew York:\nConcept learning: an Information processing\nHunt, E. B.\n1962.\nNew York: John Wiley and Sons, Inc.,\nproblem .\nand Foder, J. A. The structure of a semantic\nLanguage, 1963, 39, 170-210.\nKatz, J.\ntheory.\nKatz, J.\nulstic\n1964 .\nKelly , G\nJ.\nJ.\nand Postal,\nde sc ri ptions .\nP. M. An Integrated theory of ling-\nCambridge, Mass. The M.I.T. Press,\nThe psychology of personal constructs:\nNew York: W. W. Norton and Company, Inc.,\n1955.\nKlein, S. Automatic paraphrasing in essay format.\nVolume I.\nsp -1602/001/00 ,\nSystem Development Corporation, Santa Monica, California, 1964 .\nKlein, S. and Simmons, R. F. Syntactic dependence and the\ncomputer generation of coherent discourse. Mechanical\nTranslation, 1963, 7(2), 50-61.\n169",
          "Receipt": {
            "TextSha256": "E119BB16BB2CB83CDB663468E5EFCBD64F4F9148BE07AD74E9C87C98B78B8AC0",
            "ProcessedAtUtc": "2026-09-16T23:48:49.9608588Z",
            "ImageSha256": "51370C0844DE181AD1ED21070C50685DEDB35C24821A0E6C8A452C633B4B59CF",
            "Language": "en-US",
            "Lines": 51,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000178.ocr.txt",
            "Image": "semantic-memory-refs-000178.png"
          }
        },
        {
          "PdfPage": 179,
          "Text": "Kuno, S. K. Multiple—Patn Syntactic Analyzer. Ma thematical\nlinguistics and automatic translation, Report No. NSF-8,\n1963, Computation Laboratory of Harvard University,\n1\n1\nCambridge, Massachusetts .\nKuno, S. K. The predic tlve analyzer.\nAssociation for Computing Machinery,\nReprin ted In Hays, 1966 .\nCommunications of the\n1965, 453-462.\nLakoff, G. On the nature of syntactic Lrregularlty. Mathe-\nmatical Linguistics and Automatic Translation. Report No.\nNSF-16, 1965, computation Laboratory of Harvard University,\nCambridge, Massachusetts .\nLamb ,\nK.\nin\nLane ,\nof\nS. The sememic approach to struc tural semantics.\nIn\nA. Romney and R. D'Andrede (Eds) ,\nTranscultural Dtudies\nAmerican Anthropologist, 1964, 66 (3), Part 2.\ncogni tion .\nH. , and Schneider, B. Some discriminative properties\nsyn tac tic struc tures .\nJournal of Verbal Learning and\nVerbal Behavior, 1963, 2(5-6), 457-461.\nlindsay, R. K. Inferen tlal memory as the basis of machines\nwhich understand natural language .\nIn Feigenbaum, E. , and\nFeldman, J. (Eds.) Computers and Thought. New York:\n1963, 217-233.\nMcGraw Hill Books Co.,\nMatthews, G. H. Analysis by synthesis of sentences of natural\nProceedlngs of 1st Int. Cong. on machine trans-\nlanguages .\nlation of languages and applied language analysis, 1961,\nTeddington, England: National Physical Laboratory .\n170",
          "Receipt": {
            "TextSha256": "0FD8A9C29B2854687DD79E38E304C63538FFD7CDF08BB41FFB454A8408D3D83C",
            "ProcessedAtUtc": "2026-09-16T23:48:50.1021786Z",
            "ImageSha256": "403626CE24F40FBC186B34B9C67C693CA11C6D819E4F7C399CA9BFB1BB9D0361",
            "Language": "en-US",
            "Lines": 42,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000179.ocr.txt",
            "Image": "semantic-memory-refs-000179.png"
          }
        },
        {
          "PdfPage": 180,
          "Text": "1\n1\n1\nW. , Edwards, D. J. , Hart, T. P.\nMcCarthy, J. , Abrahams, P.\n.5 Programmer's Manual, Cambridge ,\nand Levin, M. 1. LISP\n962.\nMassachusetts:\nthe M.I.T. Computation Center,\nMiller, G. A. The magical number seven, plus or minus two:\nsome limits on our capacity for processing Informa tlon .\nPsychological Review, 1956, 63, • 81-96.\nFinite models of language\nand Chomsky, N.\nMiller, G. A. ,\nusers. In Luce, R. D. , Bush, R. L. , and Galanter, E. ,\n(Eds). Handbook of ma thematical psychology, Volume II,\n1963,\nNew York: John 'Alley and Sons, Inc.,\nMiller, G. A. ,\nand Pribram, K. H. Plans and\nGalanter, E. ,\nthe structure of behavior. New York: Holt, 1960.\nFroceedings\nMinsky, • M. Steps toward artificial intelligence.\nof the Institute of radio engineers, 1961) 49:\n8-30.\nMinsky, M. A selected descriptor-indexed bibliography to the\nliterature on artifieial intelligence.\nIn Feigenbaum, E. A. ,\nand Feldman, J. (Eds.) Computers and Thought. New York:\nMcGraw-Hill Book co.,\nNewell, A. (Ed.) IPL-V\nandum RM-3739 —RC, The\nfornia, 1963.\nShaw, J. C.\nNewell, A. ,\n1963, 453—456.\nInc . ,\nprogrammer's reference manual .\nRAND Corpora tion, Santa Monica,\nMemor -\nCali-\nand Simon, H. A. The processes of\ncreative thinking. In H. E. Gruber, G. Torrell, and M.\nWertheimer (Eds.) Contemporary approaches to creative thinking.\nNew York: Atherton Press, 1962, 63-119.\n171",
          "Receipt": {
            "TextSha256": "C1F116BCA184075FB6C561E4D512EC6AD9E2F8317A0396E44505D516E8CF8ED7",
            "ProcessedAtUtc": "2026-09-16T23:48:50.2422172Z",
            "ImageSha256": "D289ABF478428D2A84B12A14D6914ADA434E6D1E1EB59D98C5E0A37AB57D2432",
            "Language": "en-US",
            "Lines": 49,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000180.ocr.txt",
            "Image": "semantic-memory-refs-000180.png"
          }
        },
        {
          "PdfPage": 181,
          "Text": "and Simon, H. A.\nChess playing pro-\nShaw, J. C. ,\ngrams and the problem of complexity .\nResearch and Development, 1958, 2, 24\nand Simon, H. A.\nC ompu te rs\nNewell, A. ,\nI.B.M. Journal of\n320-335.\nin psychology. In\nR. D. Luce, R. Bush, and E. Galanter (Eds.), Handbook or\nNew York: John diley\nMathematical Psychology, Volume l.\nand Sons, 1963, 361-408.\nAn example of human chess play\nand Simon, H. A.\nNewell, A. ,\nPittsburgh,\nin the light of chess playing programs.\nCarnegie Institute of Technology, 1964,\nPennsylvania :\n(dittoed) .\nOgden, C . K. The General Basic English Dictionary.\nW. W. Norton and Company, Inc . ,\nNew York:\nOlney, J. Building a concept network for retrieving Informa-\ntlon from large libraries:\nYart I. Technical Memorandum\nTIM-634/001/11, System Development Corporation, Santa Monica,\nCalifornia, 1962.\nSome patterns observod in the contextual special-\nOlney, J. C.\nIzation o? word senses.\nInforma tion Storage and Retrieval,\n1964, 2, 79-101.\nOsgood, E. C. ,\nSucl, G. J. and Tannenbaum, P . H. The measure-\nment of meaning. Urbana, Illå.nois:\nUniversity of Illinois\nPress, 1957.\nOsgood, C . E. On understanding and creating sentences.\nAmerican Psy.chologist, 1965, 18, 735-751 •\n172\n1",
          "Receipt": {
            "TextSha256": "D1B33F5D1B6CBE033B3F63704C98CDAFCD564BFFA026EAAEE915579D8FD16918",
            "ProcessedAtUtc": "2026-09-16T23:48:50.4109244Z",
            "ImageSha256": "09775401C127FC20BD991CB867F796F6AFCD529D306F1C6F3025A790DA4F8953",
            "Language": "en-US",
            "Lines": 45,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000181.ocr.txt",
            "Image": "semantic-memory-refs-000181.png"
          }
        },
        {
          "PdfPage": 182,
          "Text": "Paige, J. J. , and Simon, H. A.\nCognitive processes in solving\nIn Kleinmuntz, B. (Ed.) , Eoblem\nalgebra word problems .\nSolving: Research, Me Chod and Theory. New York: john\nWiley and Sons, 1966.\nPetrick, S. R. A recogni tion procedure for transformational\ngrammars. Unpublished Ph.D. disserta tlon, M.I . T. ,\nCambridge ,\nMassachusetts, 1965.\nPiaget, J. The psychology of intelligence. Transla ted by\nM. Cook and D. E. Berlyne. London, England: Routledge and\nKegan Paul, 1950.\nQuill Ian, R. A design for an understanding machine .\nPa pe r\npresented at a colloquium:\nSemantic problems in natural\nlanguage. King's College, Cambridge University, England,\nSeptember, 1961.\n(a)\nMechanical Translation, 1962, 7, 17-29.\nQuillian, R. A seman tic coding technique for mechanical\nEnglish paraphrasing.\nInternal Memorandum of the Mechanical\nTranslation Group, Research Laboratory of Elec tronlcs,\n(b)\nM.I.T., Cambridge, Massachusetts, August, 1962.\nQuilllan, R. A notation for representing conceptual Informa-\ntlon: an application to semantics and mechanical English\nparaphrasing. SP -1395, System Development Corporation,\nSanta Monica, California, 1963.\n173",
          "Receipt": {
            "TextSha256": "E0B7FB88E4BC1AA098A49E36870F944FB44E9725F27E27C85D4361C95E0EFA0B",
            "ProcessedAtUtc": "2026-09-16T23:48:50.5416312Z",
            "ImageSha256": "7589BC7C36FB88D70F46129A214DC2FD44083D9B18776163529C86A7C33EEE42",
            "Language": "en-US",
            "Lines": 32,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000182.ocr.txt",
            "Image": "semantic-memory-refs-000182.png"
          }
        },
        {
          "PdfPage": 183,
          "Text": "Quillian, R . , Wortman, P. and Baylor, G. W. The programmable\nPiaget: behavior from the standpoint of a radical computer-\n1st. Unpublished dittoed paper, Carnetie Institute of\nTechnology, 1965.\nRaphael, B. A Computer program which\nceedings of AFIPS, 1964, Fall Joint\nPP. 577-5e9.\n'understands . t'\nPro-\nComputer Conference,\nA quantita tive study of meaning by a. con-\nRazran, G. H. S.\nditioned salivary technique (semantic condi tioning) .\nScience, 1939a, 90' 89-90 [102, 106 ] .\nReich, P. A.\nA stratifica tional theory of language acquisi-\ntion. Working Paper (IP-A), Department of Psychology\nand Mental Health Research Institute, University of Michigan,\nAnn Arbor, Michigan, 1966 .\nReid, L. S. , Henneman, R. H. and Long, E. R. An experimental\nanalysis of set:\nthe effect of categorical restriction .\nAmerican Journal of Psychology, 1960, 73, 568-572.\nReiss, R. F. An abstract machine based on classical associa-\ntlon psychology. Technical\nDivision, General Precision,\n1961.\nReitman, W. R. Cognition and\ncessing approach. New York:\n1965.\nMemorandum of Librascope\nInc.,\nGlendale, California,\nthought: an informa tion pro-\nJohn Wiley and Sons, Inc . ,\n174",
          "Receipt": {
            "TextSha256": "54DF2BE00351968C4BEEBE121C56B7DEF960B6788FC5BBBF1A31D51A290FA09C",
            "ProcessedAtUtc": "2026-09-16T23:48:50.6759348Z",
            "ImageSha256": "4E61DCE010837416F5608BA499D41B388EAD7121DE45C707DD9A3EB6327259E8",
            "Language": "en-US",
            "Lines": 36,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000183.ocr.txt",
            "Image": "semantic-memory-refs-000183.png"
          }
        },
        {
          "PdfPage": 184,
          "Text": "Problems in automatic word disambiguation .\nRubenstein, H.\nPaper presen ted at a conference on Computer-Ai ded Semantic\nResearch, held at Las Vegas, Nevada, December, 1965.\nSynthe tic language behavior. Data Processing\nSimmons, R. F.\nManagement, 1963, 5 (12) ,\n11-18.\nAn Informa tion-\nSimon, H. A. and Feigenbaum, E. A.\nprocessing theory of some effects of similarity, familiar-\nization, and meaningfulness in verbal learning. Journal\nof Verbal Learning and Verbal Behavior, 1964, 3, 385—397 •\nSimon, H. A. and Kotovsky,\nfor sequential patterns.\n5314 —546 .\nJ . Human acquisi tion of concepts\nPsychological Review, 1963, 70'\nSkinner, B. F. Verbal Behavior.\nInc., New York, 1965.\nAppleton Century Crofts,\nYngve, V . H. A model and an hypothesis for language structure.\nProceedings of the American Philosophical Society, 1960, 104,\n444-466.\nYngve, V . H. Comit Programmer's Reference Manual.\nMassachusetts: M.I.T. Press, 1961.\n175\nCambridge ,",
          "Receipt": {
            "TextSha256": "59D03F447D036F8AA3C4916608971C64DC8645507CC917FF18A8F5BAF1476174",
            "ProcessedAtUtc": "2026-09-16T23:48:50.8058684Z",
            "ImageSha256": "00D1392CD7E55647B630E5250E256987BC6DA844F2C0064E3EBEF7911021A110",
            "Language": "en-US",
            "Lines": 28,
            "Engine": "Windows.Media.Ocr",
            "Output": "semantic-memory-refs-000184.ocr.txt",
            "Image": "semantic-memory-refs-000184.png"
          }
        }
      ]
    },
    {
      "Slug": "temporal-recall",
      "Paper": "Holographic Model of Temporal Recall",
      "AtlasYear": 1968,
      "Status": "indexed",
      "Method": "publisher-reference-section",
      "SourceUrl": "https://www.nature.com/articles/217104a0",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "References",
          "PdfPages": [],
          "Text": "1. Gabor, D. Nature 161, 777 (1948).\n2. Gabor, D. Proceedings of the Royal Society A 197, 454 (1949).\n3. Julesz, B., and Pennington, K. Journal of the Optical Society of America 55, 604 (1965).\n4. Chance, B., Pye, K., and Higgins, J. J. IEEE Spectrum 4, no. 8, 79 (1967)."
        }
      ],
      "PublisherReferences": [
        {
          "key": "BF217104a0_CR1",
          "doi-asserted-by": "publisher",
          "first-page": "777",
          "DOI": "10.1038/161777a0",
          "volume": "161",
          "author": "D Gabor",
          "year": "1948",
          "unstructured": "Gabor, D., Nature, 161, 777 (1948).",
          "journal-title": "Nature"
        },
        {
          "key": "BF217104a0_CR2",
          "first-page": "454",
          "volume": "197",
          "author": "D Gabor",
          "year": "1949",
          "unstructured": "Gabor, D., Proc. Roy. Soc., A, 197, 454 (1949).",
          "journal-title": "Proc. Roy. Soc."
        },
        {
          "key": "BF217104a0_CR3",
          "first-page": "604",
          "volume": "55",
          "author": "B Julesz",
          "year": "1965",
          "unstructured": "Julesz, B., and Pennington, K., J. Opt. Soc. Amer., 55, 604 (1965).",
          "journal-title": "J. Opt. Soc. Amer."
        },
        {
          "key": "BF217104a0_CR4",
          "doi-asserted-by": "publisher",
          "first-page": "79",
          "DOI": "10.1109/MSPEC.1967.5215528",
          "volume": "4",
          "author": "B Chance",
          "year": "1967",
          "unstructured": "Chance, B., Pye, K., and Higgins, J. J., IEEE Spectrum, 4, No. 8, 79 (1967).",
          "journal-title": "IEEE Spectrum"
        }
      ],
      "Notes": [
        "All four references were checked against the complete References section on the Nature article page. Publisher-deposited Crossref metadata is retained separately. The original reference list omits article titles; no titles have been inferred."
      ],
      "ReferenceCount": 4
    },
    {
      "Slug": "theorem-proving-question-answering",
      "Paper": "The Use of Theorem-Proving Techniques in Question-Answering Systems",
      "AtlasYear": 1968,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.kestrel.edu/people/green/publications/green-raphael.pdf",
      "PdfSha256": "89939E75A3D06C32FF93D7614F7FC8A9D3D555CA069E8ADB68CEA438EEEE254F",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 413,
          "EndLine": 456,
          "PdfPages": [
            12,
            13
          ],
          "Text": "1 R F SIMMONS Answering english questions by computer: a survey Comm ACM Vol 8 No 1 January 1965\n2 K M COLBY H ENEA Heuristic methods for computer understanding of na ural language in context-restricted on-line dialogu~ Dept of Computer Sciences Stanford University 196\n\n\fThe Use of Theorem-Proving Techniques in Question-Answering Systems 181\n\n3 J A CRAIG et al DEACON: direct english access and control AFIPS Proc FJCC Vol 29 1966\n4 R E LEVIEN M E MARON A computer system for inference, execution and data retrieval Comm ACM Vol 10 No 11 pp 715-721 November 1967\n\n5 J McCARTHY Situations, actions and casual laws Memo No 2 Stanford Artificial Intelligence Project\nStanford University July 1963\n\n6 R QUILLIAN AFIPS Proc SJCC Vol 30 1967\n\n7 R F SIMMONS An approach toward answering english questions from\ntext AFIPS Proc FJCC Vol 29 1966\n\n8 J R SLAGLE Experiments with a deductive Q-A program Comm ACM Vol 8 No 12 December 1965\n\n9 F B THOMPSON English for computer AFIPS Proc FJCC Vol 29\n\n1966\n\n10 J W WEIZENBAUM ELIZA---a computer program for the study of natural language communication between man and machine Comm ACM Vol 9 No 1 January 1966\n\n11 L S COLES An on-line question-answering system with natural language and pictorial input (Paper to be presented at the ACM Conference August 1968)\n12 J L D A R L I N G T O N Machine methods for improving logical arguments expressed in english Mechanical Translation Vol 8 Nos 3 and 4 pp 41-47 June and October 1965\n13 J MC CARTHY Programs with common sense Memo No 7 Stanford Artificial Intelligence Project Stanford University September 1963\n14 B RAPHAEL A computer program which \"understands' AFIPS Proc FJCC Vol 26 1964\n\n15 B RAPHAEL SIR: A computer program for semantic information retrieval MAC-TR2 Project MAC MIT June 1964\n16 J A ROBINSON A machine-oriented logic based on the resolution principle\nJ ACM Vol 12 No 1 January 1965\n17 A NEWELL Unpublished seminar talk\n18 B RAPHAEL Aspects and applications of symbol manipulation Proc 1966 National Conference ACM 1966\n19 A NEWELL J C S H A W H A SIMON Empirical explorations of the logic theory machine: a case study in heuristics Paper presented at the Western Joint Computer Conference Los Angeles February 28 1957\n20 F BLACK A deductive question-answering system Harvard University Ph D Thesis 1964\n21 D C COOPER Theorem proving in computers Adances in Programming and Non-Numerical Computation L FOX ed Pergamon Press 1966\n22 J A ROBINSON A review of automatic theorem-proving American Mathematical Society Symposia on Applied Mathematics XIX 1967 Rice University (to be published\n23 E MENDELSON Introduction to mathematical logic van Nostrand 1964\n24 KALISH and MONTAGUE Logic: techniques of formal reasoning Harcourt Brace and World 1964\n25 L WOS et al The unit preference strategy in theorem proving AFIPS Proc FJCC Vol 26 1964\n26 T P HART A useful algebraic property of Robinson's unification algorithm Memo No 91 AI Project Project MAC MIT 1965\n27 J R SLAGLE Automatic theorem-proving with renameable and semantic resolution J ACM Vol 14 No 4 October 1967\n"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "perceptrons",
      "Paper": "Perceptrons: An Introduction to Computational Geometry",
      "AtlasYear": 1969,
      "Status": "indexed",
      "Method": "reviewed-bibliographic-facts-from-book-transcription",
      "SourceUrl": "https://www.scribd.com/document/424210282/Perceptrons",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "Bibliographic Notes: cited works, printed pages 281-287",
          "PdfPages": [],
          "Text": "Rosenblatt, F. (1959). Two theorems of statistical separability in the perceptron. Proceedings of a Symposium on the Mechanization of Thought Processes. HMSO, London, 421-456.\nRosenblatt, F. (1962). Principles of Neurodynamics. Spartan Books, New York.\nSamuel, A. L. (1959). Some studies in machine learning using the game of checkers. IBM Journal of Research and Development 3(3), 210-223.\nSamuel, A. L. (1967). Some studies in machine learning using the game of checkers, Part II. IBM Journal of Research and Development 11(4), 601-618.\nPalmieri, G., and Sanna, R. (1960). Methodos 12(48). [Title not supplied in the source.]\nGamba, A., Gamberini, L., Palmieri, G., and Sanna, R. (1961). Further experiments with PAPA. Nuovo Cimento Supplement 2, volume 20, 221-231.\nAshby, W. R. (1952). Design for a Brain. Wiley, New York. [Mentioned twice in the notes.]\nClark, W. A., and Farley, B. G. (1955). Generalization of pattern-recognition in a self-organizing system. Proceedings of the Western Joint Computer Conference, 85-111.\nMinsky, M. (1954). Neural nets and the brain-model problem. Doctoral dissertation, Princeton University. [Year as printed.]\nUttley, A. M. (1956). Conditional probability machines. Automata Studies. Princeton University, 253-285.\nAgmon, S. (1954). The relaxation method for linear inequalities. Canadian Journal of Mathematics 6(3), 382-392.\nBlock, H. D. (1962). The perceptron: a model for brain functioning. Reviews of Modern Physics 34(1), 123-135.\nPapert, S. (1961). Some Mathematical Models of Learning. Proceedings of the Fourth London Symposium on Information Theory, C. Cherry (ed.). Academic Press, New York.\nNilsson, N. (1965). Learning Machines. McGraw-Hill, New York.\nMinsky, M., and Selfridge, O. G. (1961). Learning in neural nets. Proceedings of the Fourth London Symposium on Information Theory, C. Cherry (ed.). Academic Press, New York.\nSelfridge, O. G. (1956). Pattern recognition and learning. Proceedings of the Third London Symposium on Information Theory. Academic Press, New York, 345.\nLettvin, J. Y., Maturana, H., McCulloch, W. S., and Pitts, W. (1959). What the frog's eye tells the frog's brain. Proceedings of the IRE 47, 1940-1951.\nHubel, D. H., and Wiesel, T. N. (1959). Receptive fields of single neurons in the cat's striate cortex. Journal of Physiology 148, 574-591.\nPitts, W., and McCulloch, W. S. (1947). How we know universals. Bulletin of Mathematical Biophysics 9, 127-147.\nMcCulloch, W. S. (1965). Embodiments of Mind. MIT Press, Cambridge, Massachusetts.\nMcCulloch, W. S., and Pitts, W. (1943). A logical calculus of the ideas immanent in neural nets. Bulletin of Mathematical Biophysics 5, 115-137. [Title wording and last page are as printed; the original article has a different title ending and ends on page 133.]\nMinsky, M. (1967). Computation: Finite and Infinite Machines. Prentice-Hall, Englewood Cliffs, New Jersey.\nTinbergen, N. (1951). The Study of Instinct. Oxford, New York.\nBledsoe, W. W., and Browning, I. (1959). Pattern recognition and reading by machine. Proceedings of the Eastern Joint Computer Conference, 225-232.\nRoberts, L. G. (1960). Pattern recognition with an adaptive network. IRE International Convention Record, part II, 66-70.\nRosenblatt, F. (1960). Perceptual generalization over transformation groups. Self-Organizing Systems. Pergamon Press, New York, 63-96.\nDertouzos, M. (1965). Threshold Logic: A Synthesis Approach. MIT Press, Cambridge, Massachusetts.\nMyhill, J., and Kautz, W. H. (1961). On the size of weights required for linear-input switching functions. IRE Transactions on Electronic Computers 10(2), 288-290.\nMuroga, S., and Toda, I. (1966). Lower bounds on the number of threshold functions. IEEE Transactions on Electronic Computers EC-15(5), 805-806.\nMuroga, S. (1965). Lower bounds on the number of threshold functions and a maximum weight. IEEE Transactions on Electronic Computers EC-14(2), 136-148.\nFeigenbaum, E. A., and Feldman, J. (1963). Computers and Thought. McGraw-Hill, New York.\nMinsky, M. (1968). Semantic Information Processing. MIT Press, Cambridge, Massachusetts.\nNewell, A., Shaw, J. C., and Simon, H. A. (1959). Report on a general problem-solving program. Proceedings of the International Conference on Information Processing, UNESCO House, 256-264.\nGuzman, A. (1968). Decomposition of a visual scene into bodies. Proceedings of the Fall Joint Computer Conference.\nHebb, D. O. (1949). The Organization of Behavior. Wiley, New York."
        },
        {
          "Section": "Handwritten additions from the 1972 printing",
          "PdfPages": [],
          "Text": "Block, H. D. (1970). A Review of Perceptrons. Information and Control 17. [The handwritten page range is not reliably legible in the available transcription.]\nNewell, A. (1969). A step toward the understanding of information processes. Science 165, August 1969. [The handwritten date and page range are not reliably legible in the available transcription.]\nMycielski, J. (1972). Review of Perceptrons. Bulletin of the American Mathematical Society 78(1), 12-15. DOI: 10.1090/S0002-9904-1972-12831-3.\nMinsky, M., and Papert, S. Re-View of Perceptrons. AI Memo, Artificial Intelligence Laboratory, MIT, Cambridge, Massachusetts. [Memo number appears as 293 in the transcription and remains unverified.]"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "The accessible scan is the 1988 expanded edition, which retains the original text with the authors' 1972 handwritten corrections. Bibliographic facts have been extracted from the entire Bibliographic Notes section; the surrounding critical essay is not reproduced. Later handwritten references are separated from the original references. Unclear handwritten metadata is explicitly marked and is not used to create Atlas links. This is edition-qualified coverage, not a claim that the 1969 first printing contains the later additions."
      ],
      "BibliographyEdition": "1988 expanded edition; original text with 1972 corrections",
      "ReferenceCount": 39
    },
    {
      "Slug": "relational-long-term-memory",
      "Paper": "Using Relational Operators to Structure Long-Term Memory",
      "AtlasYear": 1969,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.ijcai.org/Proceedings/69/Papers/051.pdf",
      "PdfSha256": "F02115F7EE7EEC975594239D80DF331280894ADE0B9F50683C52462DA081630C",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 391,
          "EndLine": 407,
          "PdfPages": [
            7,
            8
          ],
          "Text": "1. Green, P. F., A. K.Wolf, C. Chomsky, and K. L a u g h e r t y . \"BASEBALL: An Automatic Quest i o n Answerer.\" Proceedings of the AFIPS Conference. Fall 1961. pp. 133-144.\n2. L i n d s a y , Robert K. \" I n f e r e n t i a l Memory as the Basis of Machines which Understand N a t u r a l Language;\" in E. A. Feigenbaum and J. Feldman (eds.) Computers and Thought. New Y o r k : M c G r a w - H i l l .\n3. Raphael, Bertram. \"A Computer Program which ' U n d e r s t a n d s ' . \" Proceedings of the AFIPS Conference. F a l l 1964. pp. 577-589.\n4. Craig, J. A., S. C. Berezner, H. C. Carney, and C. R. L o n g y e a r . \"DEACON: D i r e c t E n g l i s h Access and C o n t r o l . \" Proceedings of the AFIPS Conference. F a l l 1966. pp. 365-380.\n5. V a l l e e , J . F . , K r u l e e , G.K., and Grau, A.A. \"Retrieval Formulae for Inquiry S/sterns.\" I n f o r m a t i o n Storage and R e t r i e v a l . February, 1968, pp. 13-26.\n6. Walker, D. E., \"SAFARI: An On-Line TextProcessing System,\" Proceedings of the American Documentation I n s t i t u t e , 1967. 4. pp. 144-147.\n7. Woods, W.A. \"Semantic I n t e r p r e t a t i o n of Engl i s h Questions on a Structured Data Base.\" Mathematical L i n g u i s t i c s and Automatic Translation. Harvard Computation Laborat o r y . NSF R e p o r t 1 7 . A u g u s t , 1 9 6 6 .\n8. Rosenbaum, P. S. \"A Grammar Base Q u e s t i o n answering Procedure.\" Communications of t h e ACM. O c t o b e r , 1967. p p . 6 3 0 - 6 3 5 .\n9. Hillraan, D. J . , \"Negotiation of Inquiries in an On-line Retrieval System.\" Information Storage and R e t r i e v a l . June, 1968,\n10. Jacobs, R.A., and P.S. Rosenbaum. E n g l i s h T r a n s f o r m a t i o n a l Grammar. Waltham, Massachusetts: B l a i s d e l l . 1968.\n1 1 . C h i l d c r a f t — T h e How and Why L i b r a r y . C h i c a g o : Field Enterprises Educational Corporation. 1968.\n12. Petrick, S. R. A Recognition Procedure for Transformational Grammars. Unpublished Ph.D. D i s s e r t a t i o n . M.I.T. June, 1965.\n13. P e t r i c k , S.R. A Program for Transformational\n\n-585-\n\n\fSyntactic Analysis. A i r Forca Cambridge Research Laboratories Report AFCRL-66-698, 14. Dutch, R.A. (ed.) Roget's Thesarus of English Words and Phrases. London: Longmans, Green, and Co. L t d . 1962. 15. Kats, J . J . , \"Recent Issues in Semantic Theory.\" Foundations of Language. May, 1967, p. 169. 16. Kochen, M., D.M. MacKay, M.E. Maron, M. S c r i ven, and L. Uhr. \"Computers and Comprehension'* in Manfred Kochen ( e d . ) ; The Growth of Knowledge. New Y o r k : W i l e y . 1967. pp. 230-243. 17. Harary, F . , Norman, R.Z., and C a r t w r t g h t , D. Structural Models: An Introduction to the Theory of D i r e c t e d Graphs. New Y o r k : W i l e y , 1965. 18. Sledd, J. A Short Introduction to English Grammar. Chicago: S c o t t , Foresman and Company. 1959. 19. Green, C.C. and Raphael, B. \"The Use of Theorem-Proving Techniques in QuestionAnswering Systems.\" Proceedings of 23rd ACM N a t i o n a l Conference. 1968. pp. 169182. 20. Kuhns, J.L. \"Answering Questions by Computer: A L o g i c a l S t u d y . \" RAND Memorandum RM-5428PR. December, 1967."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "teachable-language-comprehender",
      "Paper": "The teachable language comprehender: a simulation program and theory of language",
      "AtlasYear": 1969,
      "Status": "indexed",
      "Method": "full-article-transcription-crosschecked-with-publisher-records",
      "SourceUrl": "https://studylib.net/doc/8079911/the-teachable-language-comprehender--tlc-",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "Bibliography, printed pages 475-476",
          "PdfPages": [],
          "Text": "BIBLIOGRAPHY\n    ABELSON, R. P., A.ND CAI~ROLL, J. D. Computer simulation of\n    individual belief systems. Amer. Behav. Sci. 8, 9 (1965), 24-30.\n    BOBROW, D. G., MURPhy, D. L., AND TEITELMAN,W. The BBN\n    LISP system. Bolt Beranek and Newman Inc., Cambridge,\n    Mass., 1968.\n    Crm~sKv, N. Aspecls of the Theory of SynSax. M I T Press, Cambridge, Mass., 1965.\n    COLLINS, A. M., AND QUILLIAN,M. R. Retrieval time from semantic memory. Rep. No. 1692. Bolt Beranek and Newman\n    Inc., Cambridge, Mass., 1968. To appear in J. Verb. Learn.\n    Verb. Behav. (1969).\n    Communications\n    of the ACM\n    475\n    Brief Survey of Computer Languages for Symbolic and Algebraic\n    Manipulation. North-Holland Pub. Co., Amsterdam, 1968.\n    REITMAN, W. 1~. Cognition and Thought: An Information Processing Approach. Wiley, New York, 1965.\n    FEIGENBAUM, E. A., AND •ELDMAN, J. Computers and Thought.\n    McGraw-Hill, New York, 1963.\n    :KuNo, S. The predictive analyzer and a path elimination technique. Comm. ACM 8, 7 (July, 1965), 453-462.\n    LONGYEAR, C. The semantic rule. Rep. No. 67TMP-55, General\n    Electric TEMPO Project, Santa Barbara, Calif., 1967.\n    MINSKY, M. Semantic Information Processing. MIT Press, Cambridge, Mass., 1968.\n    NEWELL, A., SHAW, J. C., AND SIMON, H . A . The processes of\n    creative thinking. In Contemporary Approaches to Creative\n    Thinking, H. E. Gruber, G. Terrell, and M. Wertheimer\n    (Eds.), Atherton Press, New York, 1962, pp. 63-119.\n    PIAGET, J. ]'he Psyehology of [ntelligence. Routledge and Kegan\n    Paul, London, 1950.\n    QUILLIAN, M. R. Semantic Memory. In Semantic Information\n    Processing, M. Minsky, The MIT Press, Cambridge, Mass.,\n    1968.\n    Word concepts: a theory and simulation of some basic\n    semantic capabilities. Behav. Sci. 12 (1967), 410-430.\n    - - - - , WORTM~N, P. AND BAYLOR,G.W. The programable Piaget:\n    Behavior from the standpoint of a radical computerist. Unpublished dittoed paper, Carnegie-Melon U., Pittsburgh, Pa.,\n    1965.\n    I~.APHAEL, B., BOBROW, D. G., FEIN, L., AND YOUNG, Z.W. A\n    SIKLOSSY, L., AND SIMON, It. A. Some semantic methods for\n    language processing. Complex information processing Paper\n    129, Carnegie-Mellon U., Pittsburgh, Penn., 1968.\n    SIMMONS, R. F. Answering English questions by computer: a\n    survey. Comm. ACM 8, 1 (Jan. 1965), 53-70.\n    - - AND BURGER, Z. F. A semantic analyzer for English sentences. Rep. No. SP 2987, Syst. Develop. Corp., Santa Monica,\n    Calif., 1968.\n    TESL~R, L., ENEA, H., AND COLBY, K. M. A directed graph\n    representation for computer simulation of belief systems.\n    Dep. Comput. Sci., Stanford U., Stanford, Calif., 1967.\n    THOMPSON, F. B. The deacon project. Rep. No. 65TMP-69,\n    General Electric TEMPO Project, Santa Barbara, Calif.,\n    1965.\n    THORNE, J. P., BRATLEY, P., AND DEWAR, H. The syntactic\n    analysis of English by machine. In Machine Intelligence 3,\n    Donald Michie (Ed.), American Elsevier, New York, 1968,\n    pp. 281-309.\n    WEIZENBAUM,J. Contextual understanding by computer. Comm.\n    ACM 10, 8 (Aug. 1967), 474-480."
        }
      ],
      "PublisherReferences": [
        {
          "key": "e_1_2_1_1_1",
          "doi-asserted-by": "publisher",
          "DOI": "10.1177/000276426500800908"
        },
        {
          "key": "e_1_2_1_2_1",
          "volume-title": "The BBN LISP system",
          "author": "BOBROW D. G.",
          "year": "1968",
          "unstructured": "BOBROW , D. G. , MURP hy, D. L., AND TEITELMAN , W. The BBN LISP system . Bolt Beranek and Newman Inc., Cambridge, Mass ., 1968 . BOBROW, D. G., MURPhy, D. L., AND TEITELMAN, W. The BBN LISP system. Bolt Beranek and Newman Inc., Cambridge, Mass., 1968."
        },
        {
          "key": "e_1_2_1_3_1",
          "volume-title": "Aspecls of the Theory of Syntax",
          "author": "CHOMSKY N.",
          "year": "1965",
          "unstructured": "CHOMSKY , N. Aspecls of the Theory of Syntax . MIT Press , Cambridge, Mass ., 1965 . CHOMSKY, N. Aspecls of the Theory of Syntax. MIT Press, Cambridge, Mass., 1965."
        },
        {
          "key": "e_1_2_1_4_1",
          "volume-title": "Mass., 1968",
          "author": "COLLINS A. M.",
          "year": "1969",
          "unstructured": "COLLINS , A. M. , AND QUILLIAN , M. R. Retrieval time from semantic memory. Rep. No. 1692. Bolt Beranek and Newman Inc., Cambridge , Mass., 1968 . To appear in J. Verb. Learn. Verb. Behav. ( 1969 ). COLLINS, A. M., AND QUILLIAN, M. R. Retrieval time from semantic memory. Rep. No. 1692. Bolt Beranek and Newman Inc., Cambridge, Mass., 1968. To appear in J. Verb. Learn. Verb. Behav. (1969)."
        },
        {
          "key": "e_1_2_1_5_1",
          "volume-title": "Computers and Thought",
          "author": "FEIGENBAUM E. A.",
          "year": "1963",
          "unstructured": "FEIGENBAUM , E. A. , AND FELDMAN , J. Computers and Thought . McGraw-Hill , New York , 1963 . FEIGENBAUM, E. A., AND FELDMAN, J. Computers and Thought. McGraw-Hill, New York, 1963."
        },
        {
          "key": "e_1_2_1_6_1",
          "doi-asserted-by": "publisher",
          "DOI": "10.1145/364995.365689"
        },
        {
          "key": "e_1_2_1_7_1",
          "volume-title": "Calif.",
          "author": "LONGYEAR C.",
          "year": "1967",
          "unstructured": "LONGYEAR , C. The semantic rule. Rep. No. 67TMP-55, General Electric TEMPO Project, Santa Barbara , Calif. , 1967 . LONGYEAR, C. The semantic rule. Rep. No. 67TMP-55, General Electric TEMPO Project, Santa Barbara, Calif., 1967."
        },
        {
          "key": "e_1_2_1_8_1",
          "volume-title": "Semantic Information Processing",
          "author": "MINSKY M.",
          "year": "1968",
          "unstructured": "MINSKY , M. Semantic Information Processing . MIT Press , Cambridge, Mass ., 1968 . MINSKY, M. Semantic Information Processing. MIT Press, Cambridge, Mass., 1968."
        },
        {
          "key": "e_1_2_1_9_1",
          "first-page": "63",
          "volume-title": "Contemporary Approaches to Creative Thinking",
          "author": "NEWELL A.",
          "year": "1962",
          "unstructured": "NEWELL , A. , SHAW , J. C. , AND SIMON , H. A. The processes of creative thinking . In Contemporary Approaches to Creative Thinking , H. E. Gruber, G. Terrell, and M. Wertheimer (Eds.), Atherton Press , New York , 1962 , pp. 63 - 119 . NEWELL, A., SHAW, J. C., AND SIMON, H. A. The processes of creative thinking. In Contemporary Approaches to Creative Thinking, H. E. Gruber, G. Terrell, and M. Wertheimer (Eds.), Atherton Press, New York, 1962, pp. 63-119."
        },
        {
          "key": "e_1_2_1_10_1",
          "volume-title": "The Psyehology of Intelligence",
          "author": "PIAGET J.",
          "year": "1950",
          "unstructured": "PIAGET , J. The Psyehology of Intelligence . Routledge and Kegan Paul , London , 1950 . PIAGET, J. The Psyehology of Intelligence. Routledge and Kegan Paul, London, 1950."
        },
        {
          "key": "e_1_2_1_11_1",
          "volume-title": "Semantic Information Processing, M. Minsky",
          "author": "QUILLIAN M. R.",
          "year": "1968",
          "unstructured": "QUILLIAN , M. R. Semantic Memory . In Semantic Information Processing, M. Minsky , The MIT Press , Cambridge, Mass ., 1968 . QUILLIAN, M. R. Semantic Memory. In Semantic Information Processing, M. Minsky, The MIT Press, Cambridge, Mass., 1968."
        },
        {
          "key": "e_1_2_1_12_1",
          "doi-asserted-by": "publisher",
          "DOI": "10.1002/bs.3830120511"
        },
        {
          "key": "e_1_2_1_13_1",
          "volume-title": "The programable Piaget: Behavior from the standpoint of a radical computerist. Unpublished dittoed paper",
          "year": "1965",
          "unstructured": "Word concepts The programable Piaget: Behavior from the standpoint of a radical computerist. Unpublished dittoed paper , Carnegie-Melon U. , Pittsburgh, Pa ., 1965 . ----, WORTMAN, P. AND BAYLOR, G. W. The programable Piaget: Behavior from the standpoint of a radical computerist. Unpublished dittoed paper, Carnegie-Melon U., Pittsburgh, Pa., 1965."
        },
        {
          "key": "e_1_2_1_14_1",
          "volume-title": "North-Holland Pub. Co.",
          "author": "RAPHAEL B.",
          "year": "1968",
          "unstructured": "RAPHAEL , B. , BOBROW , D. G. , FEIN , L. , AND YOUNG , Z. W. A Brief Survey of Computer Languages for Symbolic and Algebraic Manipulation . North-Holland Pub. Co. , Amsterdam , 1968 . RAPHAEL, B., BOBROW, D. G., FEIN, L., AND YOUNG, Z. W. A Brief Survey of Computer Languages for Symbolic and Algebraic Manipulation. North-Holland Pub. Co., Amsterdam, 1968."
        },
        {
          "key": "e_1_2_1_15_1",
          "volume-title": "Cognition and Thought: An Information Processing Approach",
          "author": "REITMAN W. R.",
          "year": "1965",
          "unstructured": "REITMAN , W. R. Cognition and Thought: An Information Processing Approach . Wiley , New York , 1965 . REITMAN, W. R. Cognition and Thought: An Information Processing Approach. Wiley, New York, 1965."
        },
        {
          "key": "e_1_2_1_16_1",
          "volume-title": "Some semantic methods for language processing. Complex information processing Paper 129",
          "author": "SIKLOSSY L.",
          "year": "1968",
          "unstructured": "SIKLOSSY , L. , AND SIMON , H. A. Some semantic methods for language processing. Complex information processing Paper 129 , Carnegie-Mellon U. , Pittsburgh, Penn ., 1968 . SIKLOSSY, L., AND SIMON, H. A. Some semantic methods for language processing. Complex information processing Paper 129, Carnegie-Mellon U., Pittsburgh, Penn., 1968."
        },
        {
          "key": "e_1_2_1_17_1",
          "doi-asserted-by": "publisher",
          "DOI": "10.1145/363707.363732"
        },
        {
          "key": "e_1_2_1_18_1",
          "volume-title": "Syst. Develop",
          "author": "SIMMONS R. F",
          "year": "1968",
          "unstructured": "SIMMONS , R. F A semantic analyzer for English sentences. Rep. No. SP 2987 , Syst. Develop . Corp., Santa Monica, Calif ., 1968 . -- AND BURGER, J. F. A semantic analyzer for English sentences. Rep. No. SP 2987, Syst. Develop. Corp., Santa Monica, Calif., 1968."
        },
        {
          "key": "e_1_2_1_19_1",
          "volume-title": "Calif.",
          "author": "TESLER L.",
          "year": "1967",
          "unstructured": "TESLER , L. , ENEA , H. , AND COLBY , K. M. A directed graph representation for computer simulation of belief systems. Dep. Comput. Sci., Stanford U., Stanford , Calif. , 1967 . TESLER, L., ENEA, H., AND COLBY, K. M. A directed graph representation for computer simulation of belief systems. Dep. Comput. Sci., Stanford U., Stanford, Calif., 1967."
        },
        {
          "key": "e_1_2_1_20_1",
          "volume-title": "Calif.",
          "author": "THOMPSON F. B.",
          "year": "1965",
          "unstructured": "THOMPSON , F. B. The deacon project. Rep. No. 65TMP-69, General Electric TEMPO Project, Santa Barbara , Calif. , 1965 . THOMPSON, F. B. The deacon project. Rep. No. 65TMP-69, General Electric TEMPO Project, Santa Barbara, Calif., 1965."
        },
        {
          "key": "e_1_2_1_21_1",
          "first-page": "281",
          "volume-title": "Machine Intelligence 3",
          "author": "THORNE J. P.",
          "year": "1968",
          "unstructured": "THORNE , J. P. , BRATLEY , P. , AND DEWAR , H. The syntactic analysis of English by machine . In Machine Intelligence 3 , Donald Michie (Ed.), American Elsevier , New York , 1968 , pp. 281 - 309 . THORNE, J. P., BRATLEY, P., AND DEWAR, H. The syntactic analysis of English by machine. In Machine Intelligence 3, Donald Michie (Ed.), American Elsevier, New York, 1968, pp. 281-309."
        },
        {
          "key": "e_1_2_1_22_1",
          "doi-asserted-by": "publisher",
          "DOI": "10.1145/363534.363545"
        }
      ],
      "Notes": [
        "The complete bibliography of the original 1969 article is indexed from its full article transcription and checked against the 22 publisher-deposited references. The transcription retains scan and column-order errors: the Raphael et al. title appears earlier than its author line. The Quillian Semantic Memory reference is the 1968 chapter version. The Weizenbaum reference is a distinct 1967 article, not the 1966 ELIZA paper."
      ],
      "BibliographyEdition": "Original 1969 journal article",
      "ReferenceCount": 22
    },
    {
      "Slug": "qa3",
      "Paper": "Application of Theorem Proving to Problem Solving",
      "AtlasYear": 1969,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.ijcai.org/Proceedings/69/Papers/023.pdf",
      "PdfSha256": "691BD5C3B3DB09FD509D5372D935E3C8BC6F95090A72720FB6C91EA3AD69999F",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 778,
          "EndLine": 803,
          "PdfPages": [
            19
          ],
          "Text": "1. J. A. Robinson, \"The Present State of Mechani c a l Theorem P r o v i n g , \" a paper presented at the Fourth Systems Symposium, Cleveland, Ohio, November 19-20, 1968 (proceedings to be published).\n2. C. Green and B. Raphael, \"The Use of TheoremProving Techniques in Question-Answering Systems,\" Proc. 23rd N a t ' l . Conf. ACM, (Thompson Book Company, Washington, D.C., 1968).\n3. C. Green, \"Theorem Proving by Resolution as a Basis for Question-Answering Systems,\" Machine I n t e l l i g e n c e 4, D. Michie and B. Meltzer, Eds. (Edinburgh University Press, Edinburgh, Scotland, 1969).\n4. N. J. Nilsson, \"A Mobile Automaton: An A p p l i cation of A r t i f i c i a l Intelligence Techniques,\" a paper presented at the International Joint Conference on A r t i f i c i a l Intelligence, Washington, D.C., May 7-9, 1969 (proceedings to be published).\n5. J. McCarthy and P. Hayes, \"Some P h i l o s o p h i c a l Problems from the Standpoint of A r t i f i c i a l I n t e l l i g e n c e , \" Machine I n t e l l i g e n c e 4, D. Michie and B. M e l t z e r , Eds. (Edinburgh U n i v e r sity Press, Edinburgh, Scotland, 1969).\n6. R. J. Waldinger and R. C. T. Lee, \"PROW: A Step Toward Automatic Program W r i t i n g , \" a paper presented at the International Joint Conference on A r t i f i c i a l I n t e l l i g e n c e , Washi n g t o n , D.C., May 7-9, 1969 (proceedings to be published).\n7. L. Wos, G. A. Robinson, and D. F. Carson, \" E f f i c i e n c y and Completeness of the Set of Support Strategy in Theorem P r o v i n g , \" J.ACM, V o l . 12, No. 4, pp. 536-541 (October 1965).\n8. J. A. Robinson, \"A Machine-Oriented Logic Based on the Resolution P r i n c i p l e , \" J.ACM, V o l . 12, No. 1, pp. 23-41 (January 1965).\n9. George E r n s t , \" S u f f i c i e n t Conditions f o r the Success of GPS,\" Report No. SRC-68-17, Systems Research Center, Case Western Reserve U n i v e r s i t y , Celveland, Ohio (July 1968).\n10. A. Hormann, \"How a Computer System Can L e a r n , \" IEEE Spectrum ( J u l y 1964).\n1 1 . L. S. Coles, \"Talking With a Robot in E n g l i s h , \" paper submitted at the International Joint Conference on A r t i f i c i a l I n t e l l i g e n c e , Washi n g t o n , D.C., May 7-9, 1969 (proceedings to be published).\n\n12. John McCarthy, Paul W. Abrahams, Daniel J. Edwards, Timothy P. H a r t , and Michael I. L e v i n , LISP 1.5 Programmer's Manual (The MIT Press, Cambridge, Mass., 1962).\n13. C. Weissman, LISP 1.5 Primer (Dickenson Publ i s h i n g Company, I n c . , Belmont, C a l i f . , 1967).\n14. Lawrence Wos and George Robinson, \"Paramodul a t i o n and Set of S u p p o r t , \" summary of paper presented at the IRIA Symposium on Automatic Demonstration at V e r s a i l l e s , France, December 1 6 - 2 1 , 1968 (proceedings to be p u b l i s h e d ) .\n15. G. Robinson and L. Wos, \"Paramodulation and Theorem-Proving in First-Order Theories with E q u a l i t y , \" Machine I n t e l l i g e n c e 4, B. Meltzer and D. M i c h i e , Eds. (Edinburgh U n i v e r s i t y Press, Edinburgh, Scotland, 1969).\n16. H. Simon, \"Experiments w i t h a H e u r i s t i c Comp i l e r , \" J.ACM, V o l . 10, pp. 493-506 (October 1963).\n17. J. R. Slagle, \"Experiments with a Deductive, Question-Answering Program,\" Comm. ACM, V o l . 8, pp. 792-798 (December 1965).\n18. R. W. Floyd, \"The V e r i f y i n g Compiler,\" Computer Science Research Review, Carnegie Mellon U n i v e r s i t y (December 1967).\n19. z. Manna, \"The Correctness of Programs,\" J. Computer and Systems Sciences, V o l . 3 (1969).\n20. J. McCarthy, \"Towards a Mathematical Science of Computation,\" Proceedings ICIP (North Holland P u b l i s h i n g Company, Amsterdam, 1962).\n2 1 . B. Raphael, \"A Computer Program Which 'Unders t a n d s ' , \" Proc. FJCC, pp. 577-589 (1964).\n22. W. S. Cooper, \"Fact R e t r i e v a l and Deductive Question Answering Information Retrieval Systems,\" J.ACM, V o l . 1 1 , pp. 117-137 ( A p r i l 1964).\n23. J. A. Robinson, \"Mechanizing Higher Order L o g i c , \" Machine I n t e l l i g e n c e 4, D. Michie and B. Meltzer, Eds. (Edinburgh University Press, Edinburgh, Scotland, 1969).\n24. R. B. B a n e r j i , \"A Language f o r Pattern Recogn i t i o n , \" Pattern Recognition, V o l . 1, No. 1, pp. 63-74 (1968).\n25. R. Kowalski, \"The Case f o r Using E q u a l i t y Axioms in Automatic Demonstration,\" paper presented at the IRIA Symposium on Automatic Demonstration at V e r s a i l l e s , France, December 1 6 - 2 1 , 1968 (proceedings to be published) ."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "planner",
      "Paper": "PLANNER: A Language for Proving Theorems in Robots",
      "AtlasYear": 1969,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.ijcai.org/Proceedings/69/Papers/030.pdf",
      "PdfSha256": "9F87575EAE3D5CC7F0D77CC369EE2BE8A57EBE9531948C481E0BF978E81B5A98",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 222,
          "EndLine": 222,
          "PdfPages": [
            7
          ],
          "Text": "1 Black, F. A Deductive Question Answering System, doctoral d i s s e r t a t i o n , Harvard. 2 Green, C. C. and Raphael, B. The Use of Theorem-proving Techniques in Question-answering Systems. Proceedings of 23rd N a t i o n a l Conf. ACM. 3 Guzman, A. and Mcintosh, H. V . , Convert, Communications of ACM, Aug. 1966. 4 H e w i t t , C. , PLANNER: A Language for* Proving Theorems, A. I. memo 137, J u l y 1967. 5 McCarthy, J . ; Abrahams, P. W.; Edwards D. J . ; H a r t , T. P.; and L e v i n , Michael I. L i s p 1.5 Programmers Manual. 6 McCarthy, J. and Hayes, P., Some P h i l o s o p h i c a l Problems from the Standpoint of A r t i f i c i a l I n t e l l i g e n c e . Stanford A. I. Memo 73. 7 Newell, A . , Shaw, J. C., and Simon, H. A . , 1959. Report on a General Problem-solving Program, Proceedings of the International Conference on I n f o r m a t i o n Processing, P a r i s : UNESCO House. 8 Slagle, J. Experiments w i t h a Deductive Quest i o n - a n s w e r i n g Program, Communications of ACM December 1965."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "carps",
      "Paper": "Computer Solution of Calculus Word Problems",
      "AtlasYear": 1969,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.ijcai.org/Proceedings/69/Papers/031.pdf",
      "PdfSha256": "97DB1B873066EAE74619CF79CAFC0F180492197C42C0DAD9FFDE2A3ADAB24EA3",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 460,
          "EndLine": 482,
          "PdfPages": [
            14
          ],
          "Text": "\n1 Ayres, F . , Theory and Problems of D i f f e r e n t i a l and I n t e g r a l C a l c u l u s , Schaum P u b l i s h i n g C o . , New Y o r k , 1 9 6 4 .\n\n2 Bobrow, D.G. \" N a t u r a l Language Input f o r a Computer Problem Solving System\", Report MAC-TR-1, P r o j e c t MAC, M . I . T . , Cambridge, M a s s . , June 1964.\n\n3 C h a r n i a k , E. , \"CARPS, A Program Which Solves C a l c u l u s Word P r o b l e m s \" , Report MAC-TR-51, P r o j e c t MAC, M . I . T . , C a m b r i d g e , M a s s . , J u l y 1968.\n\n4 Guzman, A . , and M c i n t o s h , H.V., \"A M i s c e l l a n y of CONVERT P r o g r a m m i n g \" , Memorandum MAC-M-346, P r o j e c t MAC, M . I . T . , Cambridge, M a s s . , A p r i l 1967.\n\n5 Guzman, A . , Communications August 1966.\n\nand M c i n t o s h , H.V., of t h e ACM, V o l . 9,\n\n\"CONVERT\", No. 8,\n\n6 Lightstone, A . H . , Concepts of Calculus, Harper and Row, New Y o r k , 1966.\n\n7 Moses, J . , \"Symbolic MAC-TR-35, P r o j e c t MAC, December 1967.\n\nI n t e g r a t i o n \" , Report M.I.T., Cambridge, Mass.,\n\n8 Thomas, G.B., Calculus and A n a l y t i c Geometry. Addison Wesley Publishing Co., Reading, Mass., 1959.\n"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "samenlaq-ii",
      "Paper": "A Net Structure Based Relational Question Answerer: Description and Examples",
      "AtlasYear": 1969,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.ijcai.org/Proceedings/69/Papers/034.pdf",
      "PdfSha256": "B046CAD1982CA929F5FC64130F454C9912336B7AA9B013A973E45EF5A2C1E46B",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 815,
          "EndLine": 905,
          "PdfPages": [
            21
          ],
          "Text": "1. Craig, J. A., Berezner, S. C, Carney, H . C , Longyear, C . R., DEACON: Direct English Access and CONtrol. Proc. AFIPS FJCC (1966), 365-380.\n2. Elliott, R. W . , A Model for a Fact Retrieval System, unpublished Ph.D. dissertation, University of Texas, Austin, Texas, 1965.\n3. Green, C. C, Raphael, B., Research on Intelligent Question-Answering System. AFCRL-67-0370, Stanford Research I n s t i tute, Menlo Park, California, May, 1967.\n4. Levien, R., Maron, M. E., Relational Data File: A Tool for Mechanized Inference execution and Data Retrieval. RM-4793-PR, The RAND Corporation, Santa Monica, California, 1965.\n\n5.\n\nA Computer System for\n\nInference Execution and Data Retrieval.\n\nRM-5085-PR, The RAND Corporation,\n\nSanta Monica, California, Sept., 1966.\n\nAlso C.ACM 10,11 (Nov. 1967), 715-\n\n721.\n\n6. Longyear, C. R. Memory Structure in\n\nDEACON Natural Language Question-\n\nAnswering Systems. P-129, General\n\nElectric Company, TEMPO, Santa\n\nBarbara, California, 1966.\n\n7. Quillian, M. R., Semantic Memory, un­\n\npublished Ph.D. dissertation, Carnegie\n\nInstitute of Technology, Pittsburgh,\n\nPennsylvania, 1966. Also AFCRL-66-\n\n189, Bolt Beranek and Newman, I n c . ,\n\nCambridge, Massachusetts, 1966.\n\n8. Raphael, B., Semantic Information Retrie­\n\nval, unpublished Ph.D. dissertation,\n\nMassachusetts Institute of Technology,\n\nCambridge, Massachusetts, 1964. Also\n\nTR-2, Project MAC, Massachusetts\n\nInstitute of Technology, Cambridge\n\nMassachusetts, 1964.\n\n9. Shapiro, S. C, A Memory Net- Structure:\n\nPresent Implementation and a Proposed\n\nLanguage, Technical Report #53, Com­\n\nputer Sciences Department, University\n\nof Wisconsin, Madison, Wisconsin,\n\nDec. 1968.\n\n10.\n\n, Woodmansee, G. H . ,\n\nKrueger, M. W . , A Semantic Associa-\n\ntional Memory Net That Learns and\n\nAnswers Questions (SAMENLAQ). Tech­\n\nnical Report #8, Computer Sciences De­\n\npartment, University of Wisconsin,\n\nMadison, Wisconsin, Jan., 1968.\n\n1 1 . Simmons, R. F., Burger, J. F., A\n\nSemantic Analyzer for English Sentences.\n\nSP-2 987, Systems Development C o r p . ,\n\nSanta Monica, California, Jan., 1968."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "shrdlu",
      "Paper": "Procedures as a Representation for Data in a Computer Program for Understanding Natural Language",
      "AtlasYear": 1971,
      "Status": "indexed",
      "Method": "windows-ocr-with-reviewed-citation-metadata",
      "SourceUrl": "https://dspace.mit.edu/server/api/core/bitstreams/bd092d44-1552-4417-a19e-a4704c94d3ff/content",
      "PdfSha256": "8972501A3C228FE5E3D4CBAEBBCD0444173FAC7FE5E7AB83D9007F7683125725",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 1,
          "EndLine": 66,
          "PdfPages": [
            457,
            458,
            459,
            460,
            461
          ],
          "Text": "1. <Bar-Hillel 1964> Bar-Hillel, Jehoshua, LANGUAGE AND INFORMATION, Addison Wesley, 1964.\n2. <Black 1964> Black, F., \"A Deductive Question Answering System,\" in Minsky (ed.) SEMANTIC INFORMATION PROCESSING, pp. 354-402.\n3. <Bobrow 1964> Bobrow, Daniel G., \"Natural Language Input for a Computer Problem Solving System,\" in Minsky (ed.), SEMANTIC INFORMATION PROCESSING, pp. 135-215.\n4. <Bobrow 1967> Bobrow, Daniel, \"Syntactic Theory in Computer Implementations,\" in Borko (ed.), AUTOMATED LANGUAGE PROCESSING, pp. 217-252.\n5. <Bobrow 1969> Bobrow, Daniel, and J.B. Fraser, \"An Augmented State Transition Network Analysis Procedure,\" Proc. of IJCAI, 1969, pp. 557-568.\n6. <Borko 1967> Borko, Harold (ed.), AUTOMATED LANGUAGE PROCESSING, John Wiley and Sons, New York, 1967.\n7. <Charniak 1969> Charniak, Eugene, \"Computer Solution of Calculus Word Problems,\" Proc. of IJCAI, 1969, pp. 303-316.\n8. <Chomsky 1957> Chomsky, Noam, SYNTACTIC STRUCTURES, Mouton and Co., The Hague, 1957.\n9. <Chomsky 1965> Chomsky, Noam, ASPECTS OF THE THEORY OF SYNTAX, M.I.T. Press, Cambridge, Mass. 1965.\n10. <Coles 1967> Coles, L. Stephen, \"Syntax Directed Interpretation of Natural Language,\" Doctoral Dissertation, Carnegie Mellon University, 1967.\n11. <Coles 1968> Coles, L. Stephen, \"An On-Line Question-Answering System With Natural Language and Pictorial Input,\" Proc. National ACM Conference, 1968, pp. 157-167.\n12. <Craig 1966> Craig, J.A., S.C. Berezner, H.C. Carney, and C.R. Longyear, \"DEACON: Direct English Access and Control,\" Proc. FJCC 1966, pp. 365-380.\n13. <Darlington 1964> Darlington, J., \"Translating Ordinary Language into Symbolic Logic,\" Memo MAC-M-149 Project MAC, M.I.T., 1964.\n14. <Earley 1966> Earley, J.C., \"Generating a Recognizer for a BNF Grammar,\" Comp Center Paper, Carnegie Mellon Univ., 1966.\n15. <Feigenbaum 1963> Feigenbaum, Edward A., and J. Feldman, COMPUTERS AND THOUGHT, McGraw-Hill, New York, 1963.\n16. <Fodor 1964> Fodor, J.A., and J.J. Katz, (ed.) THE STRUCTURE OF LANGUAGE, Prentice Hall, Englewood Cliffs, N.J., 1964.\n17. <Fodor 1967> Fodor, J.A., and M. Garrett, \"Some Syntactic Determinants of Sentential Complexity,\" PERCEPTION AND PSYCHOPHYSICS, 1967, Vol. 2(7).\n18. <Garvin 1965> Garvin, P.L., et. al., \"A Syntactic Analyzer Study - Final Report,\" Bunker-Ramo Corp., Rome Air Development Center, RADC-TR-65-309, Dec. 1965.\n19. <Green 1969a> Green, Cordell, \"Application of Theorem Proving to Problem Solving,\" Proc. of IJCAI, 1969, pp. 219-240.\n20. <Green 1969b> Green, Cordell, and B. Raphael, \"The Use of Theorem-Proving Techniques in Question-Answering Systems,\" Proc. of ACM National Conference, 1968, pp. 169-181.\n21. <Green, P. 1961> Green, P.F., A.K. Wolf, C. Chomsky, and K. Laugherty, \"BASEBALL: An Automatic Question-Answerer,\" in Feigenbaum and Feldman (ed.) COMPUTERS AND THOUGHT, pp. 207-216.\n22. <Halliday 1961> Halliday, M.A.K., \"Categories of the Theory of Grammar,\" WORD 17, 1961.\n23. <Halliday 1966a> Halliday, M.A.K., \"Some Notes on 'Deep' Grammar,\" JOURNAL OF LINGUISTICS 2, 1966.\n24. <Halliday 1966b> Halliday, M.A.K., \"The English Verbal Group: A Specimen of a Manual of Analysis,\" Nuffield Programme in Linguistics and English Teaching, Work Paper VI, 1966.\n25. <Halliday 1967> Halliday, M.A.K., \"Notes on Transitivity and Theme in English,\" JOURNAL OF LINGUISTICS 3, 1967.\n26. <Halliday 1970> Halliday, M.A.K., \"Functional Diversity in Language as Seen From a Consideration of Modality and Mood in English,\" FOUNDATIONS OF LANGUAGE 6 (1970), pp. 322-361.\n27. <Hewitt 1969> Hewitt, Carl, \"PLANNER: A Language for Proving Theorems in Robots,\" Proc. of IJCAI, 1969, pp. 295-301.\n28. <Hewitt 1970> Hewitt, Carl, PLANNER, MAC-M-385, Project MAC, M.I.T. October, 1968, revised August, 1970.\n29. <Huddleston 1965> Huddleston, R.D., \"Rank and Depth,\" LANGUAGE 41, 1965.\n30. <Hudson 1967> Hudson, R.A., \"Constituency in a Systemic Description of the English Clause,\" LINGUA 17, 1967.\n31. <Katz 1964> Katz, J.J., and J.A. Fodor, \"The Structure of a Semantic Theory,\" in Fodor and Katz (ed.) THE STRUCTURE OF LANGUAGE, pp. 479-518.\n32. <Kellogg 1968> Kellogg, C., \"A Natural Language Compiler for On-line Data Management,\" Proc. of FJCC, 1968, pp. 473-492.\n33. <Klima 1964> Klima, Edward, \"Negation in English,\" in Fodor and Katz (ed.), THE STRUCTURE OF LANGUAGE, pp. 246-323.\n34. <Kuno 1965> Kuno, S., \"The Predictive Analyzer and a Path Elimination Technique,\" CACM 8:7 (July 1965), pp. 453-462.\n35. <Lindsay 1964> Lindsay, Robert, \"Inferential Memory as the Basis of Machines Which Understand Natural Language,\" in Feigenbaum and Feldman (ed.) COMPUTERS AND THOUGHT, pp. 217-236.\n36. <McConlogue 1965> McConlogue, K.L., and R. Simmons, \"Analyzing English Syntax with the Pattern-Learning Parser,\" CACM 8:11 (November 1965), pp. 687-698.\n37. <Michie 1968> Michie, D. (ed.), MACHINE INTELLIGENCE 3., American Elsevier Press, New York, 1968.\n38. <Miller 1951> Miller, George, LANGUAGE AND COMMUNICATION, McGraw-Hill, New York, 1951.\n39. <Minsky 1965> Minsky, Marvin, \"Matter, Mind, and Models,\" in Minsky (ed.), SEMANTIC INFORMATION PROCESSING, pp. 425-432.\n40. <Minsky 1968> Minsky, Marvin (ed.), SEMANTIC INFORMATION PROCESSING, M.I.T. Press, Cambridge, Mass., 1968.\n41. <Minsky 1970> Minsky, Marvin, \"Form and Content in Computer Science,\" JACM, Jan., 1970.\n42. <NAS 1966> National Academy of Sciences, LANGUAGE AND MACHINES: COMPUTERS IN TRANSLATION AND LINGUISTICS, National Academy of Sciences, Washington D.C., 1966.\n43. <Petrick 1965> Petrick, S., \"A Recognition Procedure for Transformational Grammars,\" Doctoral Dissertation, M.I.T., 1965.\n44. <Quillian 1966> Quillian, M. Ross, \"Semantic Memory,\" in Minsky (ed.) SEMANTIC INFORMATION PROCESSING, pp. 216-270.\n45. <Quillian 1969> Quillian, M. Ross, \"The Teachable Language Comprehender,\" CACM 12:8 (August 1969), pp. 459-475.\n46. <Raphael 1964> Raphael, Bertram, \"SIR: A Computer Program for Semantic Information Retrieval,\" in Minsky (ed.) SEMANTIC INFORMATION PROCESSING, pp. 33-134.\n47. <Robinson 1965> Robinson, J.A., \"A Machine-Oriented Logic Based on the Resolution Principle,\" JACM 12:4 (October 1965), pp. 536-541.\n48. <Shapiro 1969> Shapiro, Stuart C., and G.H. Woodmansee, \"A Net Structure Based Relational Question Answerer: Description and Examples,\" Proc. of IJCAI, 1969, pp. 325-346.\n49. <Siklossy 1968> Siklossy, L., \"Natural Language Learning by Computer,\" Doctoral Dissertation, Carnegie Mellon Univ., 1968.\n50. <Simmons 1966> Simmons, R.F., J.F. Burger, and R.E. Long, \"An Approach Toward Answering English Questions from Text,\" Proc. of FJCC 1966, pp. 357-363.\n51. <Simmons 1968> Simmons, R.F., J.F. Burger, and R. Schwarcz, \"A Computational Model of Verbal Understanding,\" Proc. of FJCC 1968, pp. 441-456.\n52. <Slagle 1965> Slagle, James R., \"Experiments with a Deductive Question-Answering Program,\" CACM, 8:12 (December 1965), pp. 792-798.\n53. <Sussman 1970> Sussman, Gerald, T. Winograd, and E. Charniak, \"Micro-Planner Reference Manual,\" AI Memo 203, Project MAC, M.I.T., July, 1970.\n54. <Tharp 1969> Tharp, Alan L., and G.K. Krulee, \"Using Relational Operators to Structure Long-Term Memory,\" Proc. of IJCAI, 1969, pp. 579-586.\n55. <Thompson 1966> Thompson, F.B., \"English for the Computer,\" Proc. of FJCC, 1966, pp. 349-356.\n56. <Thorne 1968> Thorne, J., P. Bratley, and H. Dewar, \"The Syntactic Analysis of English by Machine,\" in Michie, D. (ed.) MACHINE INTELLIGENCE 3., pp. 281-310.\n57. <Thorne 1969> Thorne, J., \"A Program for the Syntactic Analysis of English Sentences,\" CACM 12:8 (August 1969), pp. 476-480.\n58. <Weizenbaum 1966> Weizenbaum, J., \"ELIZA\" CACM 9:1 (January 1966), pp. 36-45.\n59. <Weizenbaum 1967> Weizenbaum, J., \"Contextual Understanding by Computers,\" CACM 10:8 (August 1967), pp. 474-480.\n60. <White 1970> White, Jon L., \"Interim LISP Progress Report,\" AI Memo 190, Project MAC, M.I.T., March, 1970.\n61. <Winograd 1968> Winograd, Terry, \"Linguistics and the Computer Analysis of Tonal Harmony,\" JOURNAL OF MUSIC THEORY 12:1, 1968, pp. 2-49.\n62. <Winograd 1969> Winograd, Terry, \"An Interpretive Theory of Language,\" unpublished term paper, M.I.T., 1969.\n63. <Woods 1967> Woods, William A., \"Semantics for a Question-Answering System,\" Report No. NSF-19, Aiken Computation Laboratory, Harvard Univ., Sept. 1967.\n64. <Woods 1968> Woods, W., \"Procedural Semantics for a Question-Answering Machine,\" Proc. FJCC, 1968, pp. 457-471.\n65. <Woods 1969> Woods, William A., \"Augmented Transition Networks for Natural Language Analysis,\" Report No. CS-1, Aiken Computation Laboratory, Harvard Univ., Dec. 1969.\n66. <Zwicky 1965> Zwicky, A.M., J. Friedman, B.C. Hall, and D.E. Walker, \"The MITRE Syntactic Analysis Procedure for Transformational Grammars,\" Proc. of FJCC, 1965, pp. 317-326."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Scanned bibliography pages processed with Windows OCR. The readable citation metadata was checked against the images and OCR; raw OCR is retained with page numbers and hashes. Spelling and source-publication inconsistencies may remain. Only reviewed Atlas matches become citation links."
      ],
      "OcrPages": [
        {
          "PdfPage": 457,
          "Text": "2.\n5,\n5,\n6.\n7.\n8.\n— pa%e 1&58\nBibl Jography\nBIBLIOGRAPHY\nBar-HI 1 lel, Jehoshua, LANGUAGE AND\nINFORMATION, son Wesley.\nDeductive Question Answering\nack 19610 Black, F\nin Minsky (ed,) SEMANTIC INFORMATION PROCESS\nSys teme\nNatural Language Input for\n< BObrow 196 IO Bobrow. Daniel G. e\na Computer Problem Solvrng System, t' In Mtnsky (ed.).\nSEMANTIC INFORMATION PROCESSING* PP. 135-215.\n'iSyntactjc Theory In Computer\n< Bobrow 1967> gobrow. Oeniel,\nTn Barko (ed.), AUTOMATEO LANGUAGE\nImplementations,\nPROCESSING, PP. 217-252.\n< Bobrow Bobrow. Daniel, end J 4B, Fraser nan Augmented\nProc. of\nState Transition Network Analysts Procedure,\nIJCAI, pp. 557-568.\n1967> Borko. Harold (ed.), AUTOMATED LANGUAGE\nPROCESSING, uchn Wiley and Sons, New York, 1967.\n\"Computer Solution of\nCharniak Eugene,\n<Charniak\nProc. of IJCAI, 196% pp. 303-\nCalculus Word Problems,\n316.\n(Chomsky 1957> Chomsky. Noam. SYNTACTIC STRUCTURES, Mouton\nand co., The Hague, 1957+\nChomsky, Noam. ASPECTS OF THE THEORY OF\n(Chomsky 1965)\n13.\nSYNTAX, M. I.T. Press, Cambridge, Mass..\n\"Syntax Directed\nL. Stephen.\n(Coles 1967> Coles,\nDocto ral Olssertatlon,\nInterpretation of Natural Language,\nCarnegie Mellon University, 1967.\n\"An On-L Ine Quest Ton-\n(Coles Coles, L. Stephen.\nAnswering System with Natural Language and Pictor Val\nProc. National ACM Conference, pp. 157-167.\nInput*\n(Craig Craig, J . A. , S.C. Berezner, H.C. Carney. end\n\"DEACON: 01 rect English Access and\nC, R, Longyear.\nProc. FJCC 1966, PP. 365-380.\nCONtro N'\n\"Translating Ordinary\n{Oarlington Darlington, J . ,\nLanguage into Symbolic Logic.\" Memo MAC—M-11'9 Project MAC,\nMIT",
          "Receipt": {
            "TextSha256": "F66316841B84C80D81A886ECE9FBF13039F4E79A834D58B70BA5B8496A85D9AB",
            "ProcessedAtUtc": "2026-09-16T22:50:33.1999136Z",
            "ImageSha256": "423E8A5E080E8F23ED118D5F70289161513FF123CD6F99BA3571E61D84BA4FF9",
            "Language": "en-US",
            "Lines": 62,
            "Engine": "Windows.Media.Ocr",
            "Output": "shrdlu-refs-000457.ocr.txt",
            "Image": "shrdlu-refs-000457.png"
          }
        },
        {
          "PdfPage": 458,
          "Text": "Bibl I ography\n- Page\n15.\n17.\n18.\n20.\n21.\n22.\n23.\n25.\n26.\n{Earley\nEarley, J. C . ,\n\"Generating a Recognizer for a\nBNF center Paper, Carnegie Mellon\n1966.\n(Feigenbaum 1963)\nFeigenbaum, Edward A. , end J. Feldman,\nCOMPUTERS AND THOUGHT, McGraw-Hill* New York,\n(Fodor 1964> Fodor, d . A\nand O. J, Katz, (ed.) THE\nSTRUCTURE OF LANGUAGE, Prentice Hall\nEnglewood Cl j ffs.\n1964.\n< Fodor 2967> Fodor, J . A\nand R, Gerrett, \"Some Syntactic\nDeterminants of Sentential Complexity,\" PERCEPTION AND\nPSYCHOPHYSICS* 1961* vol. 2(7).\n(Garvin 1965> Garvin, P.L., eti a 1. ,\nSyntactlc Analyzer\nStudy -\nFT nal Repo r t,\"\nBunker-namo Corp., Rome Atr\nDevelopment center, Dec. 1965.\n\"Green Green, Cordell, i'Agpl 1 cation of Theorem\nProving to Problem Sol vlng,\nProc. of IJCAI.\n1969, pp. 219-\n200.\n(Green 1969b> Green, Cordell, end 8, Rephael,\n\"The Use of\nTheorem-Proving Techni ques In Quest Jon—AnswerIng Systems.\nProc. of ACY National Conference. 1968„ pp. 169-181.\n(Green, P. Green. p . F. , A +K. Wolf* C, Chomsky, end K,\nLaugher tye\n\"BASEBALQ An Automatic Ouestlon Answe rer,\nFeigenbaum end Feldman Cedi) COMPUTERS AND THOUGHT, pp.\n207-216.\n\"Categories of the Theory\nof Grammar, n 1961.\n<Ha111day 1966e> Halliday, M.A.K\nNotes on 'Oeep\nJOURNAL OF LINGUIST lcs 1966.\n(Halliday 1966b)\n\"The Verbal\nGroup:; A Speclmen of a Manual of Nuffield\nProgrem•me in Linguist lcs and Eng I Tsh Teaching, Work Paper\n1966.\n<Hai1Tday 1967> M.A.K\nT'NOtes on Transltfvlty\nand Theme Tn Engi i sh.lr\nJOURNAL OF L INGUISTICS 1967.\n<Ha11iday 2970> M.A.K.,\n\"Functional Diversity In\nLanguage as Seen From a Consideration of Modal i ty and Mood\nIn Engl Ishi\"' FOUNDATIONS OF LANGUAGE (1910) * pp. 322-361.",
          "Receipt": {
            "TextSha256": "FE3D60E4D583F1146257177996F3C7ECBA0C978561599AE86163CEE2E8305F20",
            "ProcessedAtUtc": "2026-09-16T22:50:33.4861663Z",
            "ImageSha256": "031142945717BB8C2A594793250B41C67AFD4E15C8AD695C7C6092CC19F4A025",
            "Language": "en-US",
            "Lines": 66,
            "Engine": "Windows.Media.Ocr",
            "Output": "shrdlu-refs-000458.ocr.txt",
            "Image": "shrdlu-refs-000458.png"
          }
        },
        {
          "PdfPage": 459,
          "Text": "Bibl Tography\n- Page\n27.\n31.\n32.\n33.\n35.\n37.\n39.\n140.\nA Language for\n(Hewitt 1969> Hewitt* Carl,\nProc. of IJCAI, 1969, pp. 295-\nProving Theorems in Robots,\n301.\nPLANNER, MAC-M-386* Project\n(Hewitt Hewitt, cerif\nMAC, M. I . T i, October, 1968* revTsed August, 1970.\n\"Rank and Depth,\n(Huddleston 1965) Hudd Is ton, R .0.,\nLANGUAGE LI, 1965.\n\"Cons tt tuencY Tn a Systemic\n(Hudson 1967> Hudson, R SA. •\nLINGUA 17,\n1961.\nDescrlptlon of the English Clause,\"\nUThe Structure of a\nand J.A. Fodor,\nKatz,\n(Katz\nin Fodor and Katz (ed.) $ THE STRUCTURE OF\nSemantic Theory,\nLANGUAGE, 479-51B.\nNatu Language CompT Ter\nKellogg,\n1968, pp. 473-\nproc+ Of FdCC,\nfor On-I ine Data Management,\n\"Negation In Engl 1st\", i i\nIn\nIma 1964> Klima, Edward S. ,\nFodor and Katz (ed.), THE STRUCTURE OF LANGUAGE, pp.\n323,\n< Kuno 1965> Kuno. S, \"The Predictive Analyzer and a Peth\nCACM (July 1965). PP. 1653-062.\nEl imJnatlon Technique,\n\"Inferential Memory as the\n(Lindsay Lindsay, Robert*\nBasis of Mechlnes Which Understand NatUra1 Language,\nFeigenbaum and Feldman led.) COMPUTERS THOUGHT, PP.\n217-236,\nand R. S IrrmonS*\n< McCona ogUe 1965> K\nEng Engl Ish Syntex with a Pattern-LeernEng Parser,\nBill (November 1965), pp. 681-698.\n{Michie 1968> Michie, D. MACHINE INTELLIGENCE 3. *\nAmericen Elsevier Press, New York. 1968.\nGeorge, LANGUAGE ANO COMMUNICATI(NI,\n1951>\nNew York, 1951.\n\"Matter, MI nd, and Models,\nMinsky $ Marvin.\n(Minsky\n) SEMANTIC INFORMATION PROCESSING, pp.\ntn Minsky (ed.\n452.\nMinsky, Marvin, (ed.) SEMANTIC INFORMATION\n(Minsky 1968>\n196B.\nPROCESS ING* M.I . T. Press, Cambridge, Mass . e",
          "Receipt": {
            "TextSha256": "8500BA3545BD821115D7283DD05C44419624B35B7FF5795FE1CF248BD80E15B4",
            "ProcessedAtUtc": "2026-09-16T22:50:33.7446027Z",
            "ImageSha256": "A647C3C5B3E4861A2503ABA217B9085373115AC87403B4F84F2618F2E2D93810",
            "Language": "en-US",
            "Lines": 70,
            "Engine": "Windows.Media.Ocr",
            "Output": "shrdlu-refs-000459.ocr.txt",
            "Image": "shrdlu-refs-000459.png"
          }
        },
        {
          "PdfPage": 460,
          "Text": "44.\n46 +\n47.\n50.\n52.\n55.\nPage\n(Minsky I g \"Insky„ Marvin,\n\"Form and Content in Computer\nSc I ence.\nJACY, Jan.,\n1970.\n(NAS Nat Tonal Adademy of SC I enceS* LANGUAGE AND\nMACHINES: COMPUTERS IN TRANSLATION AND\nNational Academy of Sciences, Washington O. C. ,\n2966.\n<Petr1ck Petrick, S.\n'iA Recognition Procedure for\nTrans fo rmatlonal Gramma\nDoctoral Dissertation, M.I.T.,\nlees.\n(Quill Tan 1966> u.\n\"Semantic Memory.\nIn\nMinsky (ed.) SEMANTIC INFORMATION PROCESSING, 216-270.\n(Quill ian Qu J Ilian, M. Ross,\n\"The Teachable Language\nComprehender,\"\n{August up. 1459-475.\n< Raphael Raphael* Bert rem. \"S I R: A Computer Program\nfor Semantic Information Retrieval,\"\nIn MT nsky (ed.)\nSEMANTIC INFORVATION PROCESS pp. 33-134.\n<Robinson 1965> Robinson,\n\"A Y-achine-Oriented Logic\nBased on the Resolution Principle,\"\nOACM, (October\n1965) * 536-542.\n< Shapiro 1969> Shapl ro. Stuart C..\nand Woodmansee,\nNet Structure Based Relet tonal Quest ron Answerer:\nDescription and Exemples,\"\nProc. of IJCAI. 126% pp. 525-\n346.\n<Sikiossy Sikiossv. L. *\niiNature1 Language Learn by\nCompu\nDoctoral DissertatTon, Cernegfe Mellon\n1968.\nimmons 1966> S R , F. *\nJ.F Burger, and R.E. Long,\nApproach Toward Answer Ing Eng' iSh Questions from Text*\nProc. of Fucc PP. 357-363.\n<Simmon$ 196E> R\nBurger. and R. Sehwercz,\ntiA Computational Model of Verbel Understand rng.\" Proc. of\nFJCC* 1968* PP. 441-1656.\n(Slagle I gö5Y Slagle, James\n*'Experlments '\"fth a\nDeduct Ive Question-Answerlng Program. CAW, {Oecember\n1965). 792-798.\n(Sussman 1 g Sussman, Gere ld, T. 1K 'nograd, and E.\nCharnlak, ' 'Yi! cro-P1enner Refe r ence Manual*ti Al Memo 203*\nProject M. July, 1970,",
          "Receipt": {
            "TextSha256": "8858053215536E2A62156A43A812F360E50C35E8C8568B2B642B68092CDF24EA",
            "ProcessedAtUtc": "2026-09-16T22:50:33.9978592Z",
            "ImageSha256": "3F82BBEFBE1382A6EB7F14A7B80D32C817881EE6F06122FA7D0AD26F1D5A2BA3",
            "Language": "en-US",
            "Lines": 64,
            "Engine": "Windows.Media.Ocr",
            "Output": "shrdlu-refs-000460.ocr.txt",
            "Image": "shrdlu-refs-000460.png"
          }
        },
        {
          "PdfPage": 461,
          "Text": "Bibl tography\n- Pege\n55.\n56.\n57.\n60.\n61.\n62,\n63.\n65.\nand G.K. Kru tee,\n<Tharp Tharp, Alan L. ,\nProc.\nRelational Operators to Structure Long—Term Memory,\nof AJCAI, 190, DP, 579-536.\nEngl 15b for the Compute r, * '\n<Thompson 1966> Thompson, F\nof FUCC„ pp. 3149-356.\n\"The\nP. Bratley, and H. Dewar,\n< Thorne 1968> Tho roe,\nin Michie, O,\nSyntactic Analysts of English by Machine,\nled.) MACH INE INTELLIGENCE 3. , pg.\nA Program for the Syntactic\n< Thorne 1969) Thorne,\nCACM 120 (August 1969),\nAnalysts of Engl Ish Sentences,\npp. 476-00.\n\"ELIZA\" CACM g: 1 (January\n<Ke1zenbaum 1966> Wei zenbaum, J\n1966), pp. 36-155.\n•iContextua1 Unders tend Ing\nzenbaum 1967> Weizenbaum. J • ,\n\" CACM (August 1967), pp. 474-1480.\nby Computers,\n\"Interim LISP Progress Reporte\n(White 1970> White, Jon L • e\nAl Memo 190* ect MAC, M. I 1970.\n\"Linguistics and the\nWinograd, Terry,\n\" JOURNAL CF MUSIC\nComputer Ana lysis of Tonal Harmony.\nTHEORY 12:18 196B*\n(Wi nograd 1969> WI nograd, Terry, \"An Interpretive Theory of\nunpubl Ished term paper, M.\n1969.\nLanguage\",\niiSemantrcs for a Question-\n\"oods 1967> Woods, Wtlllam\nAnswering Sys Report No. NSF-I% Alken Computation\nsept. *1967.\nLaboratory, Harvard Uni v. ,\n<Woods 1968> Woods, W. \"Procedural Semantics for a\nProc. FJCC, pp. 1'57-411.\nQuest Ion—Answer Machine.\nAugmented Transl tron\n{Woods 1969> Woods, WI lam A\nReport CS-I,\nNetworks for Natural Language Ana\n1069.\nDec.\nAl ken Compu tat Ion Laboratory. Hervard\nO, Friedman, S.C. Hall, and\n< ZWTCky Cky,\n\"The MITRE Syntactic Anal ysls procedure for\nD.E. walker.\nof FJCC, 1965, pp. 317-\nTransformational Grammars,\n326.",
          "Receipt": {
            "TextSha256": "CD12DE92F11C6B05E0AAE79F557919D131C155D829C4B49C76E3D342B556B248",
            "ProcessedAtUtc": "2026-09-16T22:50:34.2000974Z",
            "ImageSha256": "AA233DE42F54B0A346FECA10494C6EE1D5B5491E7085AB3F930742D35CC5C869",
            "Language": "en-US",
            "Lines": 70,
            "Engine": "Windows.Media.Ocr",
            "Output": "shrdlu-refs-000461.ocr.txt",
            "Image": "shrdlu-refs-000461.png"
          }
        }
      ]
    },
    {
      "Slug": "mycin",
      "Paper": "An artificial intelligence program to advise physicians regarding antimicrobial therapy",
      "AtlasYear": 1973,
      "Status": "partial",
      "Method": "publisher-deposited-crossref-references",
      "SourceUrl": "https://api.crossref.org/works/10.1016/0010-4809(73)90029-3",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "Publisher-deposited references",
          "StartLine": 1,
          "EndLine": 26,
          "PdfPages": [],
          "Text": "1. Amarel. 1972. Medical decision making and computer modeling. Proceedings of the Fifth Hawaii International Conference on System Sciences, Supplement on Computers in Biomedicine. 173\n2. Bleich. 1971. The computer as a consultant. New. Eng J. Med.. 284. 141. 10.1056/NEJM197101212840307\n3. Bleich. 1972. Computer-based consultation—electrolyte and acid-base disorders. Amer. J. Med.. 53. 285. 10.1016/0002-9343(72)90170-2\n4. Colby. 1971. Artificial paranoia. Artificial Intelligence. 2. 1. 10.1016/0004-3702(71)90002-6\n5. Edwards. 1972. A comprehensive surveillance system of infections and antimicrobials used at Presbyterian-St. Luke's Hospital, Chicago. Amer. J. Public Health. 62. 1053. 10.2105/AJPH.62.8.1053\n6. Fox. 1970. A survey of question-answering systems. Medical Computing: Progress and Problems\n7. Gorry. 1968. Experience with a model of sequential diagnosis. Comput. Biomed. Res.. 1. 490. 10.1016/0010-4809(68)90016-5\n8. Green. 1963. BASEBALL: An automatic question-answerer (1961). Computers and Thought. 207\n9. Isner. 1972. An inferential processor for interacting with biomedical data using restricted natural language. Proceedings of AFIPS Spring Joint Computer Conference. 1107\n10. Kulikowski. 1970. Pattern recognition approach to medical diagnosis. IEEE Trans. Systems Science and Cybernetics. SSC-6. 173. 10.1109/TSSC.1970.300338\n11. Kulikowski. 1971. Computer-based models for glaucoma\n12. Kulikowski. 1972. The medical consultant program—glaucoma\n13. Lusted. 1971. Decision-making studies in patient management. New Eng. J. Med.. 284. 416. 10.1056/NEJM197102252840805\n14. Macaraeg. 1971. A study of hospital staff attitudes concerning the comparative merits of antibiotics. Clin. Pharm. Ther.. 12. 1. 10.1002/cpt19711211\n15. McCarthy. 1962\n16. Minsky. 1961. Steps toward artificial intelligence. Computers and Thought. 406\n17. Minsky. 1963. Steps toward artificial intelligence. Computers and Thought. 406\n18. Pople. 1972. An information processing approach to theory formation in biomedical research. Proceedings of AFIPS Spring Joint Computer Conference. 1125\n19. Schank. 1972. Conceptual dependency: A theory of natural language understanding. Cognitive Psychology. 3. 552. 10.1016/0010-0285(72)90022-9\n20. Sheiner. 1972. Computer-aided drug dosage. Proceedings of AFIPS Spring Joint Computer Conference. 1093\n21. Simmons. 1965. Answering English questions by computer—A survey. Commun. ACM. 8. 53. 10.1145/363707.363732\n22. Simmons. 1970. Natural language questions-answering systems. Commun. ACM. 13. 15. 10.1145/361953.361963\n23. Warner. 1964. Experience with Bayes' theorem for computer diagnosis of congenital heart disease. Ann. N.Y. Acad. Sci.. 115. 2. 10.1111/j.1749-6632.1964.tb50648.x\n24. Winograd. 1972. Understanding natural language. Cognitive Psychology. 3. 1. 10.1016/0010-0285(72)90002-3\n25. Woods. 1970. Transition network grammars for natural language analysis. Commun. ACM. 13. 591. 10.1145/355598.362773\n26. Wortman. 1972. Medical diagnosis: an information processing approach. Comput. Biomed. Res.. 5. 315. 10.1016/0010-4809(72)90065-1"
        }
      ],
      "PublisherReferences": [
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB1",
          "series-title": "Proceedings of the Fifth Hawaii International Conference on System Sciences, Supplement on Computers in Biomedicine",
          "first-page": "173",
          "article-title": "Medical decision making and computer modeling",
          "author": "Amarel",
          "year": "1972"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB2",
          "doi-asserted-by": "crossref",
          "first-page": "141",
          "DOI": "10.1056/NEJM197101212840307",
          "article-title": "The computer as a consultant",
          "volume": "284",
          "author": "Bleich",
          "year": "1971",
          "journal-title": "New. Eng J. Med."
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB3",
          "doi-asserted-by": "crossref",
          "first-page": "285",
          "DOI": "10.1016/0002-9343(72)90170-2",
          "article-title": "Computer-based consultation—electrolyte and acid-base disorders",
          "volume": "53",
          "author": "Bleich",
          "year": "1972",
          "journal-title": "Amer. J. Med."
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB4",
          "doi-asserted-by": "crossref",
          "first-page": "1",
          "DOI": "10.1016/0004-3702(71)90002-6",
          "article-title": "Artificial paranoia",
          "volume": "2",
          "author": "Colby",
          "year": "1971",
          "journal-title": "Artificial Intelligence"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB5",
          "doi-asserted-by": "crossref",
          "first-page": "1053",
          "DOI": "10.2105/AJPH.62.8.1053",
          "article-title": "A comprehensive surveillance system of infections and antimicrobials used at Presbyterian-St. Luke's Hospital, Chicago",
          "volume": "62",
          "author": "Edwards",
          "year": "1972",
          "journal-title": "Amer. J. Public Health"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB6",
          "series-title": "Medical Computing: Progress and Problems",
          "article-title": "A survey of question-answering systems",
          "author": "Fox",
          "year": "1970"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB7",
          "doi-asserted-by": "crossref",
          "first-page": "490",
          "DOI": "10.1016/0010-4809(68)90016-5",
          "article-title": "Experience with a model of sequential diagnosis",
          "volume": "1",
          "author": "Gorry",
          "year": "1968",
          "journal-title": "Comput. Biomed. Res."
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB8",
          "series-title": "Computers and Thought",
          "first-page": "207",
          "article-title": "BASEBALL: An automatic question-answerer (1961)",
          "author": "Green",
          "year": "1963"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB9",
          "series-title": "Proceedings of AFIPS Spring Joint Computer Conference",
          "first-page": "1107",
          "article-title": "An inferential processor for interacting with biomedical data using restricted natural language",
          "author": "Isner",
          "year": "1972"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB10",
          "doi-asserted-by": "crossref",
          "first-page": "173",
          "DOI": "10.1109/TSSC.1970.300338",
          "article-title": "Pattern recognition approach to medical diagnosis",
          "volume": "SSC-6",
          "author": "Kulikowski",
          "year": "1970",
          "journal-title": "IEEE Trans. Systems Science and Cybernetics"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB11",
          "article-title": "Computer-based models for glaucoma",
          "author": "Kulikowski",
          "year": "1971"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB12",
          "article-title": "The medical consultant program—glaucoma",
          "author": "Kulikowski",
          "year": "1972"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB13",
          "doi-asserted-by": "crossref",
          "first-page": "416",
          "DOI": "10.1056/NEJM197102252840805",
          "article-title": "Decision-making studies in patient management",
          "volume": "284",
          "author": "Lusted",
          "year": "1971",
          "journal-title": "New Eng. J. Med."
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB14",
          "doi-asserted-by": "crossref",
          "first-page": "1",
          "DOI": "10.1002/cpt19711211",
          "article-title": "A study of hospital staff attitudes concerning the comparative merits of antibiotics",
          "volume": "12",
          "author": "Macaraeg",
          "year": "1971",
          "journal-title": "Clin. Pharm. Ther."
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB15",
          "author": "McCarthy",
          "year": "1962"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB16_1",
          "series-title": "Computers and Thought",
          "first-page": "406",
          "article-title": "Steps toward artificial intelligence",
          "author": "Minsky",
          "year": "1961"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB16_2",
          "series-title": "Computers and Thought",
          "first-page": "406",
          "article-title": "Steps toward artificial intelligence",
          "author": "Minsky",
          "year": "1963"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB17",
          "series-title": "Proceedings of AFIPS Spring Joint Computer Conference",
          "first-page": "1125",
          "article-title": "An information processing approach to theory formation in biomedical research",
          "author": "Pople",
          "year": "1972"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB18",
          "doi-asserted-by": "crossref",
          "first-page": "552",
          "DOI": "10.1016/0010-0285(72)90022-9",
          "article-title": "Conceptual dependency: A theory of natural language understanding",
          "volume": "3",
          "author": "Schank",
          "year": "1972",
          "journal-title": "Cognitive Psychology"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB19",
          "series-title": "Proceedings of AFIPS Spring Joint Computer Conference",
          "first-page": "1093",
          "article-title": "Computer-aided drug dosage",
          "author": "Sheiner",
          "year": "1972"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB20",
          "doi-asserted-by": "crossref",
          "first-page": "53",
          "DOI": "10.1145/363707.363732",
          "article-title": "Answering English questions by computer—A survey",
          "volume": "8",
          "author": "Simmons",
          "year": "1965",
          "journal-title": "Commun. ACM"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB21",
          "doi-asserted-by": "crossref",
          "first-page": "15",
          "DOI": "10.1145/361953.361963",
          "article-title": "Natural language questions-answering systems",
          "volume": "13",
          "author": "Simmons",
          "year": "1970",
          "journal-title": "Commun. ACM"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB22",
          "doi-asserted-by": "crossref",
          "first-page": "2",
          "DOI": "10.1111/j.1749-6632.1964.tb50648.x",
          "article-title": "Experience with Bayes' theorem for computer diagnosis of congenital heart disease",
          "volume": "115",
          "author": "Warner",
          "year": "1964",
          "journal-title": "Ann. N.Y. Acad. Sci."
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB23",
          "doi-asserted-by": "crossref",
          "first-page": "1",
          "DOI": "10.1016/0010-0285(72)90002-3",
          "article-title": "Understanding natural language",
          "volume": "3",
          "author": "Winograd",
          "year": "1972",
          "journal-title": "Cognitive Psychology"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB24",
          "doi-asserted-by": "crossref",
          "first-page": "591",
          "DOI": "10.1145/355598.362773",
          "article-title": "Transition network grammars for natural language analysis",
          "volume": "13",
          "author": "Woods",
          "year": "1970",
          "journal-title": "Commun. ACM"
        },
        {
          "key": "10.1016/0010-4809(73)90029-3_BIB25",
          "doi-asserted-by": "crossref",
          "first-page": "315",
          "DOI": "10.1016/0010-4809(72)90065-1",
          "article-title": "Medical diagnosis: an information processing approach",
          "volume": "5",
          "author": "Wortman",
          "year": "1972",
          "journal-title": "Comput. Biomed. Res."
        }
      ],
      "Notes": [
        "The publisher deposits 26 Crossref records under 25 bibliography keys: reference 16 is represented by original and reprint records. All deposited records are retained, but reference 15 supplies only McCarthy and 1962, and several other records omit publication details. Attempts to retrieve the original publisher pages and Stanford HPP-73-10 did not recover the reference pages. This remains partial until the original bibliography can be checked; missing fields are not guessed."
      ]
    },
    {
      "Slug": "persistent-neural-states",
      "Paper": "The Existence of Persistent States in the Brain",
      "AtlasYear": 1974,
      "Status": "indexed",
      "Method": "publisher-reprint-reference-section",
      "SourceUrl": "https://link.springer.com/chapter/10.1007/978-1-4613-0411-1_12",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "References",
          "PdfPages": [],
          "Text": "1. C. M. Smith. The Brain. G. P. Putmanns, New York (1970).\n2. A. L. Hodgkin and A. F. Huxley. Nature 144, 710 (1939).\n3. L. I. Schiff. Quantum Mechanics, 3rd edition. McGraw-Hill, New York (1968).\n4. H. A. Kramers and G. H. Wannier. Physical Review 60, 252 (1941).\n5. G. F. Newell and E. W. Montroll. Reviews of Modern Physics 25, 353 (1953).\n6. K. Huang. Statistical Mechanics. Wiley, New York (1963).\n7. J. Ashkin and W. E. Lamb, Jr. Physical Review 64, 159 (1943).\n8. E. N. Lassettre and J. P. Howe. Journal of Chemical Physics 9, 747, 801 (1941).\n9. T. D. Lee and C. N. Yang. Physical Review 87, 410 (1952).\n10. M. R. Moldover and W. A. Little. Physical Review Letters 15, 54 (1965).\n11. J. N. Franklin. Matrix Theory. Prentice-Hall, Engelwood Cliffs, New Jersey (1968).\n12. Ibid., page 275. [Refers to Franklin, Matrix Theory, reference 11.]"
        }
      ],
      "PublisherReferences": [
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB1",
          "series-title": "The Brain",
          "author": "Smith",
          "year": "1970"
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB2",
          "doi-asserted-by": "crossref",
          "first-page": "710",
          "DOI": "10.1038/144710a0",
          "volume": "144",
          "author": "Hodgkin",
          "year": "1939",
          "journal-title": "Nature"
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB3",
          "series-title": "Quantum Mechanics",
          "author": "Schiff",
          "year": "1968"
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB4",
          "doi-asserted-by": "crossref",
          "first-page": "252",
          "DOI": "10.1103/PhysRev.60.252",
          "volume": "60",
          "author": "Kramers",
          "year": "1941",
          "journal-title": "Phys. Rev."
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB5",
          "doi-asserted-by": "crossref",
          "first-page": "353",
          "DOI": "10.1103/RevModPhys.25.353",
          "volume": "25",
          "author": "Newell",
          "year": "1953",
          "journal-title": "Rev. Mod. Phys."
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB6",
          "series-title": "Statistical Mechanics",
          "author": "Huang",
          "year": "1963"
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB7",
          "doi-asserted-by": "crossref",
          "first-page": "159",
          "DOI": "10.1103/PhysRev.64.159",
          "volume": "64",
          "author": "Ashkin",
          "year": "1943",
          "journal-title": "Phys. Rev."
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB8",
          "doi-asserted-by": "crossref",
          "first-page": "747",
          "DOI": "10.1063/1.1750835",
          "volume": "9",
          "author": "Lassettre",
          "year": "1941",
          "journal-title": "J. Chem. Phys."
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB9",
          "doi-asserted-by": "crossref",
          "first-page": "410",
          "DOI": "10.1103/PhysRev.87.410",
          "volume": "87",
          "author": "Lee",
          "year": "1952",
          "journal-title": "Phys. Rev."
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB10",
          "doi-asserted-by": "crossref",
          "first-page": "54",
          "DOI": "10.1103/PhysRevLett.15.54",
          "volume": "15",
          "author": "Moldover",
          "year": "1965",
          "journal-title": "Phys. Rev. Letters"
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB11",
          "series-title": "Matrix Theory",
          "author": "Franklin",
          "year": "1968"
        },
        {
          "key": "10.1016/0025-5564(74)90031-5_BIB12",
          "series-title": "Matrix Theory",
          "first-page": "275",
          "author": "Franklin",
          "year": "1968"
        }
      ],
      "Notes": [
        "All twelve references were recovered from the publisher's complete reference section for its reprint of the 1974 paper. They were compared with the twelve publisher-deposited records for the original journal article. Reference 12 retains the original Ibid. and identifies its antecedent."
      ],
      "BibliographyEdition": "Publisher reprint of the 1974 paper in From High-Temperature Superconductivity to Microminiature Refrigeration",
      "ReferenceCount": 12
    },
    {
      "Slug": "associative-memory",
      "Paper": "Associative Memory: A System-Theoretical Approach",
      "AtlasYear": 1977,
      "Status": "indexed",
      "Method": "publisher-back-matter-text",
      "SourceUrl": "https://link.springer.com/content/pdf/bbm%3A978-3-642-96384-1/1",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "References 1-147, printed pages 160-164",
          "PdfPages": [
            1,
            2,
            3,
            4,
            5
          ],
          "Text": "Iteferences\n1 J.R. Anderson, C.H. Bower: Human Associative Memory (Winston & Sons, Washington, D.C. 1973)\n2 D.G. Bobrow, A. Collins: Representation and Understanding (Academic Press, New\nYork 1975)\n3 J.A. Feldman, P.O. Rovner: Comm. ACM 12, 439 (1969)\n4 M. Minsky: Computation: Finite and Infinite Machines (Prentice-Hall, Englewood Cliffs 1967)\n5 D.E. Rumelhart, P.H. Lindsay, D.A. Norman: A Process Model for Long-Term Memory. In Organization and Memory, ed. by E. Tulving, W. Donaldson (Academic Press,\nNew York 1972)\n6 D.A. Savitt, H.H. Love, Jr., R.E. Troop: In 1967 Spring Joint Computer Conf.,\nAFIPS Conf. Proc. (AFIPS Press, Montvale 1967) p. 87\n7 R.C. Schank, K.M. Colby: Computer Models of Thought and Language (W.H. Freeman\n& Co., San Francisco 1973)\n8 H.A. Simon, A. Newell: Information-Processing in Computer and Man. In Perspec\u0002tives on the Computer Revolution. ed. by Z.W. Pylyshyn (Prentice-Hall,\nEnglewood Cliffs 1970)\n9 R.J. Collier: IEEE Spectrum l, 67 (1966)\n10 A.G. Hanlon: IEEE Trans. EC-~ 509 (1966)\n11 P.J. van Heerden: Appl. Opt. ~ 393 (1963)\n12 T. Kohonen: IEEE Trans. C-21, 353 (1972)\n13 G.R. Knight: Appl. Opt. lJ. 904 (1974)\n14 B. Parhami: Proc. IEEE 61. 722 (1973)\n15 M. Sakaguchi, N. Nishida, T. Nemoto: IEEE Trans. C-~. 1174 (1970) 16 K. Steinbuch: Automat und Mensch (Springer, Berlin, Heidelberg, New York 1963)\n17 G.W. Stroke: An Introduction to Coherent optics and Holography (Academic Press,\nNew York 1966) 18 D.J. Willshaw, H.C. Longuet-Higgins: Associative Memory Models. In Machine Intel\u0002Ligence, vol.5, ed. by B. Meltzer and D. Michie (Edinburgh University Press, Edin\u0002burgh, 1970)\n19 A.Albert: Regression and the Moore-Penrose Pseudoinverse (Academic Press, New\nYork 1972)\n20 A. Ben-Israel, T.N.E. Greville: GeneraLized Inverses: Theory and Applications\n(Wiley-Interscience, New York 1974)\n21 T.L. Boullion, P.L. Odell: Generalized Inverse Matrices (Wiley-Interscience, New York 1971)\n22 R.E. Cline: SIAM J. Appl. Math. 12,588 (1964)\n23 T.N.E. Greville: SIAM Rev. II, 15 (1960)\n24 T. Kohonen, E. Reuhkala, K. Makisara, L. Vainio: Biol. Cyb. 22, 159 (1976)\n25 T.O. Lewis, P.L. Odell: Estimation in Linear ModeLs (Prentice-Hall, Englewood Cliffs 1971)\n26 J.E. Mayer, M. Goeppert-Mayer: StatisticaL Mechanics (Wiley, New York 1940)\n27 R. Penrose: Proc. Cambridge Philos. Soc. 51, 406 (1955)\n28 R. Penrose: Proc. Cambridge Philos. Soc. 22, 17 (1956)\n29 C.R. Rao, S.K. Mitra: GeneraLized Inverse of Matrices and Its AppLications (Wiley- Interscience, New York 1971)\n30 D. Bobrow, B. Raphael: ACM Compo Surv. Q, 153 (1974) 161\n31 F.R.A. Hopgood: Compiling Techniques (Elsevier Publishing Co., New York 1969) 32 L.R. Johnson: Comm. ACM 4, 218 (1961)\n33 M.D. McIlroy: Comm. ACM Q, 101 (1963)\n34 R. Morris: Comm. ACM 11, 38 (1968)\nl5 W.W. Peterson: IBM J. Res. Dev. 1, 130 (1957)\n36 G. Schay, W.G. Spruth: Comm. ACM ~, 459 (1962)\n37 J.G. Williams: Comm. ACM 14, 172 (1971)\n38 C.C. Foster: IEEE Trans. C~l1, 788 (1968)\n39 J. Minker: Comput. Rev. 12, 453 (1971)\n40 M.J.E. Golay: IEEE Trans. C-18, 733 (1969)\n41 B. Gold, C.M. Rader: Digital Processing of Signals (McGraw-Hill, New York 1969)\n42 S.B. Gray: IEEE Trans. C-20, 551 (1971)\n43 T. Kohonen: Tutkimus ja Tekniikka ~, 7 (1972)\n44 T. Kohonen, E. Oja: Biol. Cyb. 21, 85 (1976)\n45 T. Kohonen, M. Ruohonen: IEEE Trans. C-22, 701 (1973)\n46 M.D. Levine: Proc. IEEE 21, 1391 (1969)\n47 T. Poggio: Biol. Cyb. 19, 201 (1975)\n48 E. RiihimKki, L.-E. HKll, T. Kohonen, P. Eistola, E. TKhti: \"Application of\nComputerized Pattern Recognition Technique to Identification of Abnormal Brain\nImages\". IV Intern. Symp. Nuclear Medicine, May 20-23, 1975, Karlovy Vary,\nCzechoslovakia\n49 H. Riittinen: M.Sc. Thesis, He-lsinki University of Technology, 1976\n50 S. Ropponen: M.Sc. Thesis, Helsinki University of Technology, 1974\n51 H. Andrews: Introduction to Mathematical Techniques in Pattern Recognition\n(Wiley, New York 1972)\n52 R.O. Duda, P.E. Hart: Pattern Classification and Scene Analysis (Wiley, New York\n1973)\n53 K.S. Fu: Sequential Methods in Pattern Recognition and Machine Learning (Academic\nPress, New York 1968)\n54 K.S. Fu: Syntactic Methods in Pattern Recognition (Academic Press, New York 1974)\n55 Y.-C. Ho, A.K. Agrawala: Proc. IEEE 56, 2101 (1968)\n56 L. Kanal: IEEE Trans. IT-20, 697 (1974)\n57 G. Nagy: Proc. IEEE 56, 836 (1968)\n58 J.J.O. Palgen: International Bibliography of Pattern Recognition in Optical and\nNon-optical Imagery (State University of New York 1969)\n59 E.A. Patrick: Fundamentals of Pattern Recognition (Prentice-Hall, Englewood Cliffs 1972)\n60 J.T. Tou, R.C. Gonzalez: Pattern Recognition Principles (Addison-Wesley, Reading\n1974)\n61 L. Uhr: Pattern Recognition, Learning, and Thought (Prentice-Hall, Englewood Cl iffs 1973)\n62 J.R. Ullman: Pattern Recognition Techniques (Butterworth, London 1973)\n63 T.Y. Young, T.W. Calvert: Classification, Estimation, and Pattern Recognition\n(Elsevier, New York 1974)\n64 S. Watanabe: Knowing and Guessing (Wiley, New York 1969)\n65 A.E. Albert, L.A. Gardner, Jr.: Stochastic Approximation and Nonlinear Regression\n(MIT Press, Cambridge, MA 1967)\n66 J.M. Mendel, K.S. Fu: Adaptive, Learning, and Pattern Recognition Systems:\nTheory and Applications (Academic Press, New York 1970)\n67 N.J. Nilsson: Learning Machines (McGraw-Hill, New York 1965)\n68 L. Schmetterer: Multidimensional Stochastic Approximation, In Multivariate\nAnalysis II, ed. by P.R. Krishnaiah (Academic Press, New York 1969)\n69 Ya.Z. Tsypkin: Adaptation and Learning in Control Systems (Academic Press,\nNew York 1971)\n70 Ya.Z. Tsypkin: Foundations of the Theory of Learning Systems (Academic Press,\nNew York 1973)\n71 M. Altman: Bull. Acad. Pol. Sci. V!, 365 (1957)\n72 O.K. Faddeev, V.N. Faddeeva: Computational Methods of Linear Algebra (W.H.\nFreeman and Co., San Francisco 1963)\n73 J.K. Hale: Ordinary Differential Equations (Wiley, New York 1969)\n74 T. Kohonen, E. Oja, M. Ruohonen: Adaptation of a Linear System to a Finite Set\nof Patterns Occurring in an Arbitrarily Varying Order, Acta Polytechnica Scandinavica, Mathematics and Computer Science Series No. 25 (1974) 162\n75 E. Oja: Lic. Techn. Thesis. Helsinki University of Technology. 1975\n76 L. Pyle: Numer. Math. 10. 86 (1967)\n77 W.T. Reid: Riaaati Differential Equations (Academic Press. New York 1972)\n78 T. Kohonen: IEEE Trans. C-Zl. 444 (1974)\n79 T. Kohonen: ~oa. 1974 Intern. Symp. MUltiple-Valued Logia, May 29-31 (West Virginia University) p. 493\n80 L. Pyle: J. ACM 11. 422 (1964)\n81 H.B. Barlow: Nature 258. 199 (1975)\n82 C. Blakemore. D.E. Mitchell: Nature 241. 467 (1973)\n83 J.P. Cavanagh: Ph.D. Thesis. Carnegie-Mellon University. 1972\n84 P.T. Chopping: Nature 217. 781 (1968)\n85 J.C. Eccles: The Physiology Of Synapses (Springer. Berlin. Heidelberg. New\nYork 1964)\n86 J.C. Eccles: In Brain and Human Behavior, ed. by A.G. Karcmar. J.C. Eccles\n(Springer. Berlin. Heidelberg. New York 1972)\n87 D. Gabor: IBM J. Res. Dev. 13. 156 (1969)\n88 J.S. Griffith: Nature (Lond.) 211. 1160 (1966)\n89 D. Hebb: ~ganiBation of Behavior (Wiley. New York 1949)\n90 P.J. van Heerden: The Foundation of Empiriaal Knowledge with a Theory of\nArtifiaial Intelligenae (Wistik. Wassenaar. Netherlands 1968) 91 G. Horn. S.P.R. Rose. P.P.G. Bateson: Science 181. 506 (1973)\n92 D.H. Hubel. T.N. Wiesel: J. Compo Neurol. 158.307 (1974)\n93 D.H. Hubel. T.N. Wiesel: J. Neurophysiol. 18.229 (1965)\n94 D.H. Hubel. T.N. Wiesel: J. Physiol. 160. 106 (1962)\n95 M.O. Huttunen: Persp. Biol. Med. 17. 103 (1973)\n96 H. Hyden. E. Egyhazi: Proc. Nat. Acad. Sci. 48. 1366 (1962)\n97 T. Kohonen. P. Lehtio. J. Rovamo: Ann. Acad. Sci. Fenn. A.V. Med. l§] (1974)\n98 Y.S. Lashley: In The Neurophysiology of Lashley; Seleated Papers of K.S.\nLashley, ed. by F.A. Beach et al. (McGraw-Hill. New York 1960)\n99 A.L. Leiman. C.N. Christian: \"Electrophysiological Analysis of Learning and\n100\n101\n102\n103\n104\n105\n106\n107\n108\n109\n110\n111\n112\n113\n114\n115\n116\n117\n118\n119\nMemory\". In The Physiologiaal Basis of Memory, ed. by J.A. Deutsch (Academic\nPress. New York 1973)\nH.C. Longuet-Higgins: Nature 217. 104 (1968)\nR. Mark: Memory and Nerve Cell Conneations. Critiaisms and Contributions from\nDevelopmental Neurophysiology (Clarendon Press. Oxford 1974)\nW.C. McCulloch. W.A. Pitts: Bull. Math. Biophysiol. ~. 115 (1943)\nV.B. Mountcastle: J. Neurophysiol. 20.408 (1957)\nG.N. Polyakov: Osnovyi sistematiki neironov novoi koryi bolahovo mozga ~eloveka\n(Medicina. Moscow 1973)\nK. Pribram: Languages of the Brain (Prentice-Hall. Englewood Cliffs 1971)\nM.R. Rosenzweig. E.L. Bennett (eds.): Neural Meahanisms of Learning and Memory (The MIT Press. Cambridge 1976)\nJ. Rovamo. J. Hyvarinen: A Physiologiaal Model of Assoaiative Memory. Experimental Brain Research (in press)\nG.M. Shepherd: The Synaptia ~ganization of the Brain (Oxford University Press.\nNew York 1974)\nG.S. Stent: Proc. Nat. Acad. Sci. USA 70. 997 (1973)\nJ. Szentagothai: Brain Research 95. 475 (1975)\nR.F. Thompson: Introduation to Physiologiaal Psyahology (Harper & Row. New\nYork 1975)\nG. Ungar: Int. J. Neurosci. 1. 193 (1972)\nL. Wall~e. J.K.S. Jansen. K. Nygaard: Kybernetik ~ 130 (1969)\nP.R. Westlake: Kybernetikl. 129 (1970)\nS.I. Amari: IEEE Trans. C-21. 1197 (1972)\nJ.A. Anderson: Math. Biosci.14. 197 (1972)\nV. Braitenberg: In Physias and Mathematias of the Nervous Sifstem, ed. by M.\nConrad. et al. (Springer. Berlin. Heidelberg. New York 1974)\nV. Braitenberg: J. Theor. Biol. i5. 421 (1974)\nL.N. Cooper: \"A Possible Organization of Animal Memory and Learning\". In ~oa.\nNobel Symp. Colleative ~operties of Physiaal Systems, ed. by B. Lundquist. S. Lundquist (Academic Press. New York 1974) 120 O.D. Creutzfeldt, U. Kuhnt, L.A. Benveneto: Exp. Brain Res. 21, 251 (1974) 121 B. Farley, W. Clark: IRE Trans. IT-!, 76 (1954)\n122 K. Fukushima: Kybernetik 12, 58 (1973)\n123 P.C. Gilbert: Brain Res. 70, 1 (1974)\n124 P. Gilbert: Nature 254, 688 (1975)\n125 S. Grossberg: Kybernetik 10, 49 (1972)\n126 R. Hess, K. Negishi, O. Creutzfeldt: EXp. Brain Res. 22, 415 (1975)\n127 T. Kohonen: A Class of Randomly Organized Assoaiative Memories, Acta Poly- technica Scandinavica, Electrical Engineering Series No. El 25 (1971)\n128 T. Kohonen: Introduation of the Prinaiple of Virtual Images in Assoaiative\nMemories, Acta Polytechnica Scandinavica, Electrical Engineering Series No.\nEl f2 (1971) 129 T. Kohonen: Int. J. Neurosci. ~, 27 (1973)\n130 H.C. Longuet-Higgins, D.J. Willshaw, O.P. Buneman: Quart. Rev. Biophys. l,\n223 (1970)\n131 C. v. d. Malsburg: Kybernetik 14, 85 (1973)\n132 D. Marr: J. Physiol, (Lond.) 202,437 (1969)\n163\n133 K. Nakano, J. Nagumo: In Advanae Papers of the Conferenae, 2nd Intern. Joint\nConf. Artifiaial Intelligenae (The British Computer Society, London 1971) p. 101\n134 K. Nakano: IEEE Trans. SCM-£, 380 (1972)\n135 M.M. Nass, L.N. Cooper: Biol. Cyb. 19, 1 (1975)\n136 F. Ratliff: Maah Bands (Holden-Day, San Francisco 1965)\n137 F. Rosenblatt: Psychol. Rev. 65, 386 (1958)\n138 F. Rosenblatt: Prinaiples of Neurodynamias: Peraeptrons and the Theory of\nBrain Meahanisms (Spartan Books, Washington, D.C. 1961)\n139 J.W. Silverstein: Biol. Cyb. 22,73 (1976)\n140 G.J. Simmons: In 1964 Spring Joint Computer Conf., AFIPS Conf. Proa. (Spartan Books, Washington, D.C. 1964) Vol. 25, p. 493\n141 A.M. Uttley: \"Conditional Probability Computing in a Nervous System\", In\nMeahanization of Thought Proaesses (H.M. Stationery Office, London 1950)\n142 B. Widrow: \"Generalization and Information Storage in Networks of Adaline\nNeurons\", In Self Organizing Systems 1962, ed, by G.T. Yovits et al. (Spartan Books, Washington, d.c. 1962)\n143 H. Wigstrom: Kybernetik 12, 204 (1973)\n144 H. Wigstrom: Kybernetik 16, 103 (1974)\n145 D. Willshaw: Ph.D. Thesis, University of Edinburgh, 1971\n146 D.J. Willshaw, O.P. Buneman, H.C. Longuet-Higgins: Nature 22Z, 960 (1969) 147 D.J. Willshaw, C. v. d. Malsburg: Proc. Roy. Soc. (London) B 194, 431 (1976)"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "The complete numbered reference list, 1 through 147, was recovered from the publisher's 1977 back matter. The following author index is excluded. PDF extraction places some reference numbers on separate lines; the original text and numbering are preserved rather than inferred from adjacent lines."
      ],
      "BibliographyEdition": "1977 first edition, DOI 10.1007/978-3-642-96384-1",
      "ReferenceCount": 147
    },
    {
      "Slug": "content-addressable-memories",
      "Paper": "Content-Addressable Memories",
      "AtlasYear": 1980,
      "Status": "indexed",
      "Method": "publisher-back-matter-text",
      "SourceUrl": "https://link.springer.com/content/pdf/bbm%3A978-3-642-96552-4/1",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "References, printed pages 331-352",
          "PdfPages": [
            1,
            2,
            3,
            4,
            5,
            6,
            7,
            8,
            9,
            10,
            11,
            12,
            13,
            14,
            15,
            16,
            17,
            18,
            19,
            20,
            21,
            22
          ],
          "Text": "References\n1.1 T. Kohonen: Associative Memory - A System-Theoretical Approach, Com\u0002munication and Cybernetics, Vol. 17 (Springer, Berlin, Heidelberg,\nNew York 1978)\n1.2 J.R. Anderson, C.H. Bower: Human Associative Memory (Winston & Sons,\nWashington, D.C. 1973)\n1.3 A.C. Hanlon: IEEE Trans. EC-15, 509-521 (1966)\n1.4 J. ~linker: Comput. Rev. 12, 453-504 (1971)\n1.5 B. Parhami: Proc. IEEE 61, 722-730 (1973)\n1.6 V. Bush: Atl. Mon. 176, 101 (1945)\n1.7 F. Rosenblatt: Frinciples of Neurodynamics: Perceptrons and the Theory\nof Brain Mechanisms (Spartan Books, Washington, D.C. 1961)\n1.8 M.H. Lewin: RCA Rev. 23, 215-229 (1962)\n1.9 V.L. Newhouse, R.E. Fruin: Electronics 35, 31-36 (1962)\n1.10 A.E. Slade, H.O. McMahon: Proc. EJCC 10, 115-120 (1956)\n1.11 J.D. Noe: Curro Res. Devel. Sci. Doc. 9 (1956)\n1.12 E.H. Frei, J. Goldberg: IEEE Trans. EC-l0, 718-722 (1961)\n1.13 A.D. Falkoff: J. ACM 9, 488-511 (1962)\n1.14 D.A. Savitt, H.H. Love, R.E. Troop: \"Association Storing Processor\",\nVol. I (AD 818 529), Vol. II (AD 818 530) (Hughes Aircraft Co. 1967)\n1.15 D.A. Savitt, H.H. Love, R.E. Troop: AFIPS Proc. SJCC 31, 87 (1967)\n1.16 J.A. Feldman, P.D. Rovner: Commun. ACM 12, 439-449 (1969)\n1.17 P.D. Rovner, J.A. Feldman: In Information Frocessing 68 (North Holland,\nAmsterdam 1969) pp. 579-585\n1.18 J. McCarthy, P.W. Abrahams, D.J. Edwards, T.P. Hart, M.I. Levin: LISP\n1.5 Frogrammer's Manual (MIT Press, Cambridge, Mass. 1962)\n1.19 H.A. Simon, A. Newell: \"Information-Processing in Computer and Man\".\nA Sigma Xi-RESA National Lecture 1964. In Perspectives on the Computer\nRevolution, ed. by Z.W. Pylyshyn (Prentice-Hall, Englewood Cliffs 1970)\n1.20 P.D. Rovner: LEAP User's Manual (MIT Lincoln Laboratory, Mass. 1968)\n1.21 K. Van Lehn: SAIL User Manual, Stanford Computer Science Report\nSTAN-CS-73-373 (1973)\n1.22 L.A. Zadeh: Inf. Control 8, 338 (1965)\n1.23 B. Widrow: \"Generalization of Information Storage in Networks of Adal ine\nNeurons\", in Self Organizing Systems 1962, ed. by G.T. Yovits (Spartan\nBooks, Washington, D.C. 1962)\n1.24 K. Steinbuch: Automat und Mensch (Springer, Berlin, Heidelberg, New York\n1963)\n1.25 P.J. van Heerden: The Foundation of Empirical KnOWledge with a Theory\nof Artificial Intelligence (Wistik, Wassenaar, Netherlands 1968)\n1.26 D. Gabor: IBM J. Res. Dev. 13, 156 (1969)\n1.27 M.D. Levine: Proc. IEEE 57, 1391 (1969)\n1.28 D. Marr: J. Physiol. London 202, 437 (1969)\n1.29 D.J. Willshaw, H.C. Longuet-Higgins: \"Associative Memory Models\", in\nMachine Intelligence, Vol. 5, ed. by B. Meltzer and D. Michie (Edinburgh\nUniversity Press, Edinburgh 1970) 332\n1.30\n1.31\n1.32\n1.33\n1.34\n1.35\n1.36\n1. 37\n1.38\n1.39\n1.40\n1.41\n1.42\n1.43\n1.44\n1.45\n1.46\n1.47\n1.48\n1. 49\n1.50\n1.51\n1. 52\n1.53\n1. 54\n1.55\n1. 56\n1.57\n1.58\n2.1\n2.2\n2.3\n2.4\n2.5\n2.6\nK. Nakano, J. Nagumo: In Advance Papers of the Conference, 2nd Intern.\nJoint Conf. Artificial Intelligence (The British Computer Society, London 1971)\nT. Kohonen: A class of randomly organized associative memories. Acta\nPoly tech. Scand., Electr. Eng. Ser. No El 29 (1971)\nJ.A. Anderson: Math. Biosci. 14, 197 (1972)\nT. Kohonen: IEEE Trans. C-21, 353 (1972)\nT. Kohonen, M. Ruohonen: IEEE Trans. C-22, 701 (1973)\nL.N. Cooper: \"A Possible Organization of Animal ~lemory and Learning\". in Proc. Nobel Symp. Collective Properties of Physical Systems, ed. by\nB. Lundquist, S. Lundquist (Academic Press, New York 1974)\nM.L. Yevick: Pattern Recognition 7, 197 (1975)\nR. Sorabji: Aristotle on Memory (Brown University Press, Providence,\nR.1. 1972)\nR.W. Hamming: Bell Syst. Tech. J. 29, 147 (1950)\nD.J. Rogers, T.T. Tanimoto: Science 132 (1960)\nL. Ornstein: J. M. Sinai Hosp. 32, 437 (1965)\nK. Sparck-Jones: Automatic Keyword Classification and Information\nRetrieval (Butterworth's, London 1971)\nJ. Minker, E. Peltola, G.A. Wilson: Tech. Report 201, University of\nMaryland, Computer Science Center (1972)\nJ.S. Lienard, M. Mlouka, J.J. Mariani, J. Sapaly: \"Real-Time Segmentation of Speech\", in Preprints of the Speech Communication Seminar, Vol. 3\n(Almqvist & Wiksell, Uppsala 1975) p. 183\nT. Tanimoto: An Elementary Mathematical Theory of Classification and\nPrediction (IBM Corp., 1958)\nJ. tukasiewicz: Ruch Filos. 5, 169 (1920)\nE.L. Post: Am. J. Math. 43, 163 (1921)\nL.A. Zadeh: IEEE Trans. SMC-3, 28 (1973)\nL.A. Zadeh, K.S. Fu, K. Tanaka, M. Shimura (eds.): FUzzy Sets and Their\nApplications to Cognitive and Decision Processes (Academic Press,\nNew York 1975)\nV.M. Velichko, N.G. Zagoruyko: Int. J. Man Mach. Stud. 2, 223 (1970)\nK. Abe: Technical Report of the Professional Group on Pattern\nRecognition of IECEJ, PRL 74-5 (1974) (In Japanese) V.I. Levenshtein: Sov. Phys. Dokl. 10, 707 (1966)\nT. Okuda, E. Tanaka, T. Kasai: IEEE Trans. C-25, 172 (1976)\nK.S. Fu: Syntactic Methods in Pattern Recognition (Academic Press,\nNew York 1974)\nK.S. Lashley: In The Neurophysiology of Lashley; Selected Papers of\nK.S. Lashley, ed. by F.A. Beach (McGraw-Hill, New York 1960)\nT. Poggio: Biol. Cybern. 19, 201 (1975)\nA.E. Albert, L.A. Gardner, Jr.: Stochastic Approximation and Nonlinear\nRegression (MIT Press, Cambridge, Mass. 1967)\nM.L. Minsky: Computation: Finite and Infinite Machines (Prentice-Hall,\nEnglewood Cliffs, N.J. 1967)\nG. Bohn: Biol. Cybern. 29, 193 (1978)\nD. Knuth: The Art of Computer Programming. Vol. 3: Sorting and Searching\n(Addison-Wesley, Reading, Mass. 1973)\nJ. Martin: Computer Data-Base Organization. 2nd printing (Prentice\u0002Hall, Englewood Cliffs, N.J. 1977)\nI. Flores: Data Structure and Management (Prentice-Hall, Englewood Cliffs, N.J. 1970)\nW.H. Desmonde: Real-Time Data Processing Systems (Prentice-Hall, London 1964)\nD. Lefkovitz: File Structures for On-Line Systems (Hayden, New York\n1969)\nC.W. Bachman: Commun. ACM 15, 628-634 (1972) 2.7 N. Chapin: Proc. FJCC 1969, pp. 413-422\n2.8 G.G. Dodd: Comput. Surv. 1, 117-139 (1969)\n2.9 R.F. Schubert: Datamation (July 1972) pp. 42-47\n2.10 D.K. Chow: Inf. Control 15, 377-396 (1969)\n2.11 S.P. Ghosh: Inf. Sci. 1,363-380 (1969)\n2.12 S.P. Ghosh, M.E. Senko: J. ACM 16, 569-579 (1969)\n2.13 S.P. Ghosh: Commun. ACM 15, 802-808 (1972)\n2.14 S.P. Ghosh: Inf. Sci. 6, 1-9 (1973)\n2.15 S.P. Ghosh: Inf. Control 25, 145-169 (1974)\n2.16 H. Hellerman: Commun. ACM 5, 205-207 (1962)\n2.17 V.Y. Lum: Commun. ACM 13, 660-665 (1970)\n2.18 D.R. Morrison: J. ACM 15, 514-534 (1968)\n2.19 E. Wong, T.C. Chiang: Commun. ACM 14, 593-597 (1971)\n2.20 CODASYL Data Base Task Group: April 1971 report (Amsterdam 1971)\n2.21 CODASYL Systems Committee: Feature Analysis of Generalized Data\nManagement Systems (Amsterdam 1971)\n2.22 R. Morris: Commun. ACM 11, 38-44 (1968)\n2.23 A. Gill: Linear Sequential Circuits: Analysis, Synthesis, and\nApplications (McGraw-Hill, New York 1966)\n2.24 T. Kohonen: Digital Circuits and Devices (Prentice-Hall, Englewood Cliffs, N.J. 1972)\n333\n2.25 A.D. Lin: Proc. AFIPS 1963 SJCC 24 (Spartan Books, New York) pp. 355-366\n2.26 M. Hanan, F.P. Palermo: IBM J. Res. Dev. 7, 127 (1963)\n2.27 G. Schay, N. Raver: IBM J. Res. Dev. 7, 121 (1963)\n2.28 E.R. Berlekamp: Algebraic Coding Theory (McGraw-Hill, New York 1968)\n2.29 V.Y. Lum, P.S.T. Yuen, M. Dodd: Commun. ACM 14, 228-239 (1971)\n2.30 R. Sprugnoli: Commun. ACM 20,841-850 (1977)\n2.31 D. Mitra: Inf. Control 23, 205-220 (1973)\n2.32 B.H. Bloom: Commun. ACM 13, 422-426 (1970)\n2.33 W.D. Maurer: Commun. ACM 11, 35-38 (1968)\n2.34 J.D. Beyer: Commun. ACM 11, 378 (1968)\n2.35 C.E. Radke: Commun. ACM 13, 103-107 (1970)\n2.36 A.C. Day: Commun. ACM 13, 481-482 (1970)\n2.37 F.R.A. Hopgood, J. Davenport: Comput. J. 15, 314-315 (1972)\n2.38 A.F. Ackerman: Commun. ACM 17, 164 (1974)\n2.39 V. Batagelj: Commun. ACM 18, 216-217 (1975)\n2.40 A. Ecker: Comput. J. 17, 340-343 (1974)\n2.41 J.R. Bell: Commun. ACM 13, 107-109 (1970)\n2.42 L. Lamport: Commun. ACM 13, 573-574 (1970)\n2.43 J.R. Bell, C.H. Kaman: Commun. ACM 13, 675-677 (1970)\n2.44 F. Luccio: Commun. ACM 15, 1045-1047 (1972)\n2.45 A.J.D. Pawson: Comput. J. 16, 285 (1973)\n2.46 S.K. Bandyopadhyay: Commun. A01 20, 262-263 (1977)\n2.47 W.D. t~aurer: Commun. ACM 11, 378 (1968)\n2.48 R.P. Brent: Commun. ACM 16, 105-109 (1973)\n2.49 C. Halatsis, G. Philokyprou: Commun. ACM 21, 554-557 (1978)\n2.50 J.A. Feldman, J.R. Low: Commun. ACM 16, 703 (1973)\n2.51 F.R.A. Hopgood: Compo Bull. 11,297-300 (1968)\n2.52 F.R.A. Hopgood: Compiling Techniques (Elsevier Publishing Co., New York\n1969)\n2.53 C. Bays: Comput. J. 16, 126-131 (1973)\n2.54 C. Bays: Commun. ACM 16, 11-14 (1973)\n2.55 D. Severance, R. Duhne: Commun. ACM 19, 314-326 (1976)\n2.56 K. Furukawa: Inf. Process. Japan 13, 13-18 (1973)\n2.57 O. Amble, D.E. Knuth: Comput. J. 17, 135-142 (1974)\n2.58 D.G. Bobrow: Commun. ACM 18, 413-415 (1975)\n2.59 L.D. Higgins, F.J. Smith: Comput. J. 14, 249-253 (1971)\n2.60 W.W. Peterson: IBM J. Res. Dev. 1, 130 (1957) 334\n2.61 W. Buchholz: IBM Syst. J. 2, 86 (1963)\n2.62 M.H. McKinney: Proc. Nat. Comput. Conf., 1977, pp. 371-377\n2.63 J.A. Feldman, P.D. Rovner: Stanford Artificial Intelligence Project\nMemo AI-66 (1968)\n2.64 J.A. Feldman, P.D. Rovner: Commun. ACM 12, 439-449 (1969)\n2.65 P.D. Rovner, J.A. Feldman: MIT Lincoln Laboratory, Technical Note\n1967-19 (1967)\n2.66 P.D. Rovner, J.A. Feldman: In Information Processing 68 (North-Holland,\nAmsterdam 1969) pp. 579-585\n2.67 M.D. McIlroy: Commun. ACM 6, 101 (1963-)\n2.68 A.P. Ershov: Doklady Akad. Nauk SSSR 118, 427-430 (1958)\n2.69 G. Schay, Jr., W.G. Spruth: Commun. ACM 5, 459-462 (1962)\n2.70 J. Kral: Comput. J. 14, 145-149 (1971)\n2.71 L.R. Johnson: Commun. ACr~ 4,218-222 (1961)\n2.72 W.P. Heising: IBM Syst. J. 2, 112 (1963)\n2.73 C.A. Olson: Proc. 1969 ACM Nat. Conf., San Francisco, pp. 539-549\n2.74 J.A. van der Pool: IBM J. Res. Dev. 16, 579 (1972)\n2.75 J.A. van der Pool: IBM J. Res. Dev. 17,27 (1973)\n2.76 M. Tainiter: J. ACM 10, 307-315 (1963)\n2.77 J.G. Williams: Commun. ACM 14, 172-175 (1971)\n2.78 V.Y. Lum, P.S.T. Yuen: Commun. ACM 15, 996-997 (1972)\n2.79 V.Y. Lum: Commun. ACM 16, 603-612 (1973)\n2.80 J.D. Ullman: J. ACM 19, 569-575 (1972)\n2.81 D. Knuth: The Art of Computer Programming, Vol. 1: Fundamental\nAlgorithms (Addison-Wesley, Reading, r'lass. 1968)\n2.82 B.F. Cheydleur: \"Dimension: An Associative Memory\", Philco Computer\nDivision, Dec. 1962\n2.83 B.F. Cheydleur: Vistas in Information Handling, Vol.1 (Spartan Books,\nWashington, D.C. 1963) p. 55\n2.84 B.F. Cheydleur: Am. Doc. 14, 56 (1963)\n2.85 P. Weston, S.M. Taylor: Coordinated Sci. Lab. Rept. R-393, Sept.\n1968 (AD-679 948)\n2.86 P.C. Patton: Computer 3, 19 (1970)\n2.87 N.S. Prywes, H.J. Gray, W.I. Landauer, D. Lefkowitz. S. Litwin:\nU. Pennsylvania, The Moore School of El. Eng., Tech. Rept. No.1\n(AD-270 573) (1961)\n2.88 N.S. Prywes, H.J. Gray: In Information Processing 1962 (North-Holland\n1962) pp. 273-278\n2.89 N.S. Prywes, H.J. Gray: AlEE Special Publication S 136 (1962)\npp. 87-101\n2.90 N.S. Prywes, H.J. Gray: IEEE Trans. Commun. and Electron. 82, 488-492\n(1963)\n2.91 N.S. Prywes: Proc. IEEE 54, 1788-1966 (1966)\n2.92 R.L. Rivest: SIAM J. Comput. 5, 19-50 (1976)\n2.93 R.A. Gustafson: Proc. Symp. on Information and Retrieval, ACM, New York,\nApril, 1971, pp. 163-174\n2.94 T. Kohonen: Associative Memory - A System-Theoretical Approach,\nCommunication and Cybernetics, Vol.17 (Springer, Berlin, Heidelberg,\nNew York 1978)\n2.95 T. Kohonen, E. Reuhkala: Helsinki University of Technology, Report TKK-F-A335, 1978\n2.96 T. Kohonen, E. Reuhkala: Proc. 4th Intern. Joint Conf. on Pattern\nRecognition, Kyoto (1978) pp. 807-809\n2.97 G. Dewey: Relative Frequency of English Speech Sounds (Harvard\nUniversity Press, Cambridge, MA 1923)\n2.98 A.G. Debus, Ed.: World Who's Who in Science, 1st ed. (Marquis Who's\nWho Inc., Chicago 1968) 2.99\n2.100\n2.101\n2.102\n2.103\n2.104\n2.105\n2.106\n2.107\n2.108\n2.109\n2.110\n2.111\n2.112\n2.113\n2.114\n2.115\n2.116\n2.117\n2.118\n2.119\n2.120\n2.121\n2.122\n2.123\n2.124\n2.125\n2.126\n2.127\n2.128\n2.129\n2.130\n2.131\n2.132\n2.133\n2.134\n2.135\n2.136\n2.137\n2.138\n2.139\n2.140\n2.141\n2.142\n2.143\n2.144\nM.C. Harrison: Commun. ACM 14, 777-779 (1971)\nC.E. Goble: Comput. J. 18, 18-20 (1975)\nA.V. Aho, M.J. Corasick: Commun. ACM 18, 333-340 (1975)\nC.N. Alberga: Commun. ACM 10, 302-313 (1967)\nF.J. Damerau: Commun. ACM 7, 171-176 (1964)\nW. Doster: IEEE Trans. Comput. C-26, 1090-1101 (1977)\nV.I. Levenshtein: Sov. Phys. Dokl. 10, 707-710 (1966)\nH.L. Morgan: Commun. ACM 13, 90-94 (1970)\nT. Okuda, E. Tanaka, T. Kasai: IEEE Trans. C-25, 172-178\n(1976)\nE.M. Riseman, A.R. Hanson: IEEE Trans. C-23, 480-493 (1974)\nJ.R. Ullman: Comput. J. 20, 141-147 (1977)\nE. Fredkin: Commun. ACM 3, 490-499 (1960)\nK. Maly: Commun. ACM 19, 409-415 (1976)\nW.A. Burkhard: J. Compo Syst. Sci. 15, 280-299 (1977)\nE.G. Coffman, Jr., J. Eve: Commun. ACM 13, 427-436 (1970)\nD.G. Severance: Comput. Surv. 6, 175-194 (1974)\nE.G. Mallach: Comput. J. 20, 137-140 (1977)\nA.I. Durney: Comput. and Autom. 5, 6-9 (1956)\nW.O. Maurer, T.G. Lewis: Computing Surv. 7, 5-19 (1975)\nP.G. Sorenson, J.P. Tremblay, R.F. Deutscher: Inf. Sci. Can. 16, 1\n(1978)\nG.D. Knott: Comput. J. 18, 265 (1975)\nA. Bookstein: J. Amer. Soc. Inf. Sci. 25, 232 (1974)\nW. Doster: Wiss. Ber. A. 51, 104 (1978)\nW.B. Samson, R.H. Davis: Comput. J. 21, 210 (1977)\nD.G. Severance, J.V. Carlis: Minnesota Univ. Man. Inf. Syst. Res.\nCenter Rept. MISRC-TR-77-06 (1977)\nT. Yuba, M. Hoshi: Trans. Inst. Electron. Commun. Eng. (Japan) E61,\n52 (1978)\nR.L. Rivest: J. ACM 25, 200 (1978)\nW. Littwin: CR Acad. Sci. Ser. A (France) 286, A695 (1978)\nW. Littwin: In Proc. 4th Int. Conf. on Very Large Data Bases (IEEE, New York 1978)\nM. Ajtai, J. Komlos, E. Szemeredi: Inf. Proc. Lett. 7, 270 (1978)\nL.J. Guibas: J. ACM 25, 544 (1978)\nL.J. Guibas, E. Szemeredi: J. Compo Syst. Sci. 16,226 (1978)\nG. Lyon: Proc. COMPCON Fall '78, p. 378 (1978)\nA. Kraml i, J. Pergel: Probl. Control Inf. Theo·ry 6, 207 (1977)\nS. Fortune, J. Hopcroft: Inf. Proc. Lett. 8, 20 (1979)\nA.L. Tharp: Inf. Syst. 4, 55 (1979)\nP.-A. Larson: BIT (Sweden) 18, 184 (1978)\nR. Fagin, N.J. Pippenger, H.R. Strong: IBM Tech. Discl. Bull. 21,\n809 (1972)\n335\nW.A. Burkhard: Proc. 11th Hawaii Int. Conf. on Syst. Sci. (1978) p. 99\nR.C. Lee, Y.H. Chin, S.C. Chang: IEEE Trans. SE-2, 185 (1976)\nE. Hill, Jr.: \"A Comparative Study of Very Large Data Bases\", in\nLecture Notes on Computer Science, Vol. 59 (Springer, Berlin, Heidel\u0002berg, New York 1978)\nM.L. Griss: Proc. 10th Hawaii Int. Conf. on Syst. Sci. (1977) p. 169\nM. Imai, T. Fukumura, Y. Yoshida: Inf. Process. Soc. Jpn. (Joho Shori) 18, 639 (1977)\nW.T. Wipke, S. Krishnan, G.I. Ouchi: J. Chern. Inf. Comput. Sci. 18,\n32 (1978)\nL. Hodes, A. Feldman: J. Chern. Inf. Comput. Sci. 18,96 (1978)\nT.G. Lewis: \"A Hashed-Array Database Technique for Minicomputers\",\nin Microprocessors, Microprogramming, and Minicomputers, ed. by\nR. Chattergy, U.W. Pooch (Western Periodicals, N. Hollywood 1977) 336\n2.145 S. Nachmens, S. Berild: Data (Sweden), No.6, 41 (1976)\n2.146 P. Heckel: Commun. ACM 21, 264 (1978)\n2.147 L.H. Groner, A.L. Goel: \"Concurrency in Hashed Fi le Access\", in\nInformation Processing 74 (IFIP, North-Holland, Amsterdam 1974)\n2.148 E. Goto, T. Ida, T. Gunji: Inf. Proc. Lett. 6,8 (1977)\n2.149 T. Ida, E. Goto: \"Performance of a Parallel Hash Hardware with Key\nDeletion\", in Information Processing 77 (IFIP, North-Holland,\nAmsterdam 1977)\n2.150 R.S. Fabry: Commun. ACM 17, 403 (1974)\n3.1 R.F. Rosin: Proc. AFIPS 1962 SJCC, p. 203\n3.2 C.C. Foster: IEEE Trans. C-17, 788 (1968)\n3.3 G.A. Anderson: IEEE Trans. C-23, 1317 (1974)\n3.4 J.T. Koo: IEEE Trans. SC-5, 208 (1970)\n3.5 F.A. Behnke: U.S. Patent No. 3,195,109, July 13 (1965)\n3.6 W. Hilberg: IEEE Trans. EC-15, 117 (1966)\n3.7 W. Hilberg: U.S. Patent No. 3,706,078, Dec. 12 (1972)\n3.8 V.F. Rudakov, Yu.I.Il 'yashenko: Star 5, 1420 (1967)\n3.9 A.V. Campi, B.H. Gray: U.S. Patent No. 3,634,829, Jan. 11 (1972)\n3.10 E.E. Davidson: IBM Tech. Discl. Bull. 17, 855 (1974)\n3.11 E.H. Frei, J. Goldberg: IRE Trans. EC-10, 718 (1961)\n3.12 J. Favor: Goodyear Aerospace Corp., AP-111770, Oct. 1964\n3.13 C.A. Hill: Goodyear Aerospace Corp., GER-12181, May 1965\n3.14 C.C. Foster, F. Stockton: IEEE Trans. C-20, 1580 (1971)\n3.15 Y. Chu: IEEE Trans. EC-14, 600 (1965)\n3.16 H.S. Stone: Proc. AFIPS 1968 FJCC, p. 949 (1968)\n3.17 K.E. Batcher: U.S. Patent No. 3,800,289, March (1974)\n3.18 K.E. Batcher: Proc. 1974 Nat. Comput. Conf., p. 405\n3.19 M.H. Lewin: RCA Rev. 23, 215 (1962)\n3.20 R.R. Seeber, A.B. Lindqvist: IBM J. Res. Dev. 6, 126 (1962)\n3.21 H. Weinstein: IEEE Trans. EC-12, 564 (1963)\n3.22 L. Johnson, M. McAndrew: IBM J. Res. Dev. 8, 189 (1964)\n3.23 H.S. Miller: IEEE Trans. EC-13, 614 (1964)\n3.24 V. Chlouba: Inf. Process. Mach. 13, 139 (1967)\n3.25 R.R. Seeber: Proc. EJCC 18, 179 (1960)\n3.26 A. Wolinsky: Commun. ACM 11, 488 (1968)\n3.27 C.V. Ramamoorthy, J.L. Turner, B.W. Bah: IEEE Trans. C-27, 800 (1978)\n3.28 K.E. Batcher: Proc. AFIPS 1968 SJCC, p. 307 (1968)\n3.29 E.J. Gauss: J. ACM 8, 418 (1961)\n3.30 J.L. Anderson: IBM Tech. Discl. Bull. 4,28 (1962)\n3.31 G. Estrin, R.H. Fuller: Proc. IEEE Pacific Comput. Conf., p. 118 (1963)\n3.32 A. Kaplan: Proc. FJCC 24, 193 (1963)\n3.33 A. Wolinsky: IEEE Symp. on Search Memory, May 1964 (IEEE New York\n1964 )\n3.34 A. Wolinsky: IEEE Trans. C-18, 899 (1969)\n3.35 A.B. Lindqvist: IBM Tech. Discl. Bull., August 1965, p. 372\n3.36 R.M. Bird, J.L. Cass, R.H. Fuller: Rome Air Dev. Center, Rept. No.\nTR-66-209, Sept. 1966, Vol 1, AD-800 387\n3.37 R.M. Bird, S.L. Cabs, R.H. Fuller: Rome Air Dev. Center, Rept. No.\nTR-66-209, Sept. 1966, Vol. 2, AD-376 572\n3.38 R.R. Seeber, A.B. Lindqvist: U.S. Patent No. 3,430,205, Feb. (1969)\n3.39 T.Y. Feng: Proc. 4th Annu. Princeton Conf. on Inf. Sci. and Syst., Princeton U., March 1970, p. 442\n3.40 C.C. Foster: Univ. Massachusetts, Comput. Inf. Sci. Dept., Tech. Note\nCS-00016, July 1970\n3.41 D.P. Agrawal: Proc. 1974 Conf. on Comput. Syst. and Technol.,\nOct. 29-Nov. 1, 1974, p. 180\n3.42 W.A. Crofut, M.R. Sottile: IEEE Trans. EC-15, 529 (1966) 337\n3.43 D.C. Alexander, R.H. Dennard, F.L. Post: IBM Advanced Systems, 17.022,\nMay 1961\n3.44 F.H. Young: Oregon State Univ., Dept. ~lath., In-House Doc. 1962\n3.45 P.T. Rux: Oregon State Univ., July 1967 (AD-660 792)\n3.46 P.T. Rux: Oregon State Univ., Feb. 1968 (AD-671 910)\n3.47 P.T. Rux: IEEE Trans. C-18, 512 (1969)\n3.48 P.T. Rux, F.W. Weingarten, F.H. Young: IEEE Comput. Group Repository,\nNo. 67-72, March 1967\n3.49 B. Parhami: Tech. Report UCLA-ENG-7213, Univ. California LA (1972)\n3.50 B. Parhami: Proc. AFIPS 1972 FJCC, p. 681 (1972)\n3.51 G.L. Hollander: Proc. 1956 JCC, p. 128 (1956)\n3.52 D. Warren: IEEE Symp. on Search Memory, May 1964 (IEEE.New York 1964)\n3.53 R.I. Roth: U.S. Patent No. 3,257,646, June 21 (1966)\n3.54 D.L. Slotnick: Adv. in Comput. 10, 291 (1970\n3.55 L.D. Healy, G.J. Lipovski, K.L. Doty: Proc. AFIPS 1972 FJCC, p. 691\n(1972 )\n3.56 N. Minsky: Proc. AFIPS 1972 FJCC, p. 587 (1972)\n3.57 G.B. Houston, R.H. Simonsen: U.S. Patent No. 3,906,455, Sep. 16 (1975)\n3.58 Chyuan Shiun Lin, D.C.P. Smith: ACM Trans. Database Syst. 1, 53 (1976)\n3.59 M. Flinders, P.L. Gardner, R.J. Llewelyn, J.F. Minshull: Proc.1970 IEEE\nInt. Comput. Group Conf., p. 314\n3.60 P.L. Gardner: IEEE Trans. C-20, 764 (1971)\n3.61 P.L. Gardner: U.K. Patent No.1 281 387, July 12 (1972)\n3.62 T. Kohonen: Digital Circuits and Devices (Prentice-Hall, Englewood Cl iffs, N.J. 1972)\n3.63 J.R. Brown, Jr.: Proc. Spec. Tech. Conf. on Nonlinear Magnetics, Los Angeles, Cal., Nov. 1961\n3.64 M.H. Lewin, H.R. Beelitz, J.A. Rajchman: Proc. AFIPS 1963 FJCC 24,\n101 (1963)\n3.65 G.G. Pick, D.B. Brick: Am. Doc. Inst., 26th Annu. Meeting, p. 245, Oct. 1963\n3.66 G.G. Pick: Proc. AFIPS 1964 FJCC 26, 107 (1964)\n3.67 E.L. Younker, C.H. Heckler, D.P. Masher, J.M. Yarborough: Stanford\nRes. Inst., Oct. 1964 (AD-609 126)\n3.68 E.L. Younker, C.H. Heckler, D.P. Masher, J.M. Yarborough: Proc. SJCC\n25, 515 (1964)\n3.69 S.T.C.: B.P. 1013241, Dec. 1965\n3.70 M.H. Lewin: U.S. Patent No. 3,245,052, Apr. 5 (1966)\n3.71 R.A. Henle, I.T. Ho, G.A. Maley, R. Waxman: Proc. FJCC 1969, p. 61 (1969)\n3.72 D.C. Wyland: Comput. Des. No.9, p. 61 (1971)\n3.73 K.E. Iverson: A Programming Language (Wiley, New York 1962)\n3.74 A.D. Falkoff: J. ACM 9, 488 (1962)\n3.75 G. Estrin: Proc. WJCC, May 1960, p. 33\n3.76 J. Ausley: Moore School of Electr. Eng., M. Sc. Thesis, 1961\n3.77 M.J. Flynn: Purdue Univ., Ph. D. Thesis, BTP-62-1782, Jun 1961\n3.78 R.H. Fuller: Disser. Absts. 24, 1960 (1963)\n3.79 R.H. Fuller: UCLA, Dept. of Eng., Rept. No. 63-25 (1963)\n3.80 Computer Command and Control: Rept. No. 5-101-5, Jan. 1964\n3.81 J.E. McAteer, J.A. Capobianco, R.L. Koppel: Proc. 1964 FJCC, p. 81\n(1964)\n3.82 A.E. Slade: Am. Doc. Inst. 27th Annu. Meeting (1964)\n3.83 S. Sohara: IEEE Symp. on Search Memory (1964)\n3.84 A.V. Campi, R.M. Dunn, B.H. Gray: IEEE Trans. AES-1, 168 (1965)\n3.85 W.F. Chow: Univac, Quart. Prog. Rept., Oct. 1965 (AD-477 446)\n3.86 W.F. Chow: Sperry Rand Corp. Quart. Rept. 1965 (AD-472 571)\n3.87 W.F. Chow: Sperry Rand Corp. Quart. Rept. 1966 (AD-804 628)\n3.88 R.W. Haas, E.H. Blevis: Marquardt Corp., July 1965 (AD-620 915)\n3.89 R.R. Seeber, A.B. Lindqvist: Proc. IFIP Congo 2, 479 (1965) 338\n3.90\n3.91\n3.92\n3.93\n3.94\n3.95\n3.96\n3.97\n3.98\n3.99\n3.100\n3.101\n3.102\n3.103\n3.104\n3.105\n3.106\n3.107\n3.108\n3.109\n3.110\n3.111\n3.112\n3.113\n3.114\n3.115\n3.116\n3.117\n3.118\n3.119\n3.120\n3.121\n3.122\n4.1\n4.2\n4.3\n4.4\n4.5\n4.6\n4.7\n4.8\n4.9\n4.10\n4.11\nR. Haas, E. Blewis, S. Requa, I. Hanlet: Marquardt Corp., AF 30(602)-\n3709, June 1966 (AD-488 453)\nR.W. Haas, J.M. Hanlet: Marquardt Corp., Final Rept. Dec. 1967\n(AD-825 274)\nR.C.M. Barnes, I.N. Hooton: H.M. Stationary Office, England 1966\nC.F. Chong, P.A. Rivelli, J.S. Mathias, P. Geaneotes: Sperry Rand\nCorp. Quart. Rept., June 1967 (AD-815 774L)\nJ.P. Bartlett: Proc. IEEE Int. Comput. Group Conf., Wash.\nJune 16-18, 1970, p. 299\nJ. Bartlett, J. Mudge, J. Springer: Electronics 43, 96 (1970)\nJ. Cashera: Air Force Avionics Lab. Rept. No. AFAL-TR-70-71,\nAug. 1970 (AD-876 691/7SL)\nElectron 13, 53 (1972)\nA.B.E. Ellis: The Marconi Review xxxv, 42 (1972)\nJ. Minker: Uni'J. of Maryland, Tech. Rept. TR-195, July 1972\nH.-O. Leil ich: \"Assoziative Speicher\", in Taschenbuch der Informatik,\nVol.I (Springer, Berlin, Heidelberg, New York 1974) p. 479\nH.-O. Leilich: \"Access Methods and Associative Memories\", Proc.\nDigitale Speicher (NTG-Fachber. Germany) 58, 328 (1977)\nG. Wolf: Elektron. Rechenanlagen 17, 264 (1975)\nG. Wolf: Data Report 11, 29 (1976)\nW. Motsch: Electron. Rechenanlagen 19, 274 (1977)\nW. Motsch: Proc. Digitale Speicher, Stuttgart, Germany, 22-24 March\n1977 (NTG-Fachber. Germany) 58, 339 (1977)\nElectronics 51, 63 (1978)\nE.S. Lee: Proc. AFIPS 1963 SJCC 23, 381 (1963)\nE.G. Wagner, J. McCarthy: U.S. Patent No. 3,093,814, June (1963)\nR.D. Ross: IBM Tech. Discl. Bull. No. 10, p.561 (1967)\nA.M. Peskin: Proc. FJCC 1970, p. 615 (1970)\nD.E. Davis, C.W. Hannaford: IBM Tech. Discl. Bull. 15, 719 (1972)\nP.A. DiGleria, M.H. Hallett: U.K. Patent No. 1,265,645, March 1 (1972)\nR.B. Derickson: Comput. Des. 7, 60 (1968)\nJ.D. Erwin, LD. Jensen: Proc. AFIP·S 1970 FJCC, p. 621 (1970)\nW.K. King: IEEE Trans. C-20, 671 (1971)\nA. Weinberger: Comput. Des. 10, 77 (1971)\nA. Weinberger: U.K. Patent No. 1 280 753, July 5 (1972)\nM. Handrich: Informatique et Gestion, No. 49, p. 95 (1973)\nN.K. Natarajan, P.A.V. Thomas: IEEE Trans. C-18, 424 (1969)\nB. Parhami: Asilomar Conf. on Circuits, Syst. and Comput.,\n7th Annu. Conf. Rec. Pap., p. 439 (1973)\nJ.F. Minshull, A.S. Murphy: U.K. Patent No.1 289 249, Sep. 3 (1972)\nD.W. Digby: IEEE Trans. C-22, 768 (1973)\nJ. Bartlett, J. Mudge, J. Springer: Electronics 43, 96 (1970)\nD. Aspinall, D.J. Kinniment, D.B.G. Edwards: Information Processing 68\n(North-Holland, Amsterdam 1969) p. 800\nD.J. Kinniment, A.E. Knowles, D.B.G. Edwards: Proc. Int. Conf. on\nMicroelectronics (IEEE, London 1969) p. 37\nD.W. Hillis, T.W. Hart: U.K. Patent No.1 255 206, Dec. 1 (1971)\nJ.B. Hughes: U.S. Patent No. 3,704,456, Nov. 28 (1972)\nSERT J. 8, 78 (1974)\nMotorola Semiconductor Products, Inc.: In Developments in Large-Scale\nIntegration (Engineering Edition). (Phoenix, Ariz. 1968)\nS. Matsue: U.S. Patent No. 3,703,709, Nov. 21 (1972)\nR.W. Murphy: IBM Tech. Discl. Bull. 14, 1669 (1971)\nH.H. Berger, C.L. Schuenemann, S.K. Wiedmann: IBM Tech. Discl. BUll.\n14, 3088 (1972)\nA.W. Bidwell, W.O. Pricer: Proc. Int. Solid-State Circuits Conf.\n(1967) p. 78 4.12 D.P. Repchick: IBM Discl. Bull. 10, 502 (1967)\n4.13 W. Hilberg: U.S. Patent No. 3,706,078, Sep. 8 (1971)\n4.14 W.O. Pricer: IBM Tech. Discl. Bull 16, 3160 (1974)\n4.15 A.A. orlikovskii: Sov. Microelectron. (USA) 6, 355 (1977)\n4.16 J.R. Burns, J.H. Scott: Proc. AFIPS 1969 FJCC, p. 469 (1969)\n4.17 L.D. Wald: Proc. 1970 NAECON, p. 277\n4.18 R.M. Lea: Electron. Lett. 8, 391 (1972)\n339\n4.19 C.J. Shead: GEC J. Sci. Tech. 40, 119 (1973)\n4.20 I.V. Prangishvili, V.A. Dementujev, M.S. Sonin: Mikroelektronika (USSR)\n4, 497 (1975)\n4.21 R. Igarashi, T. Kurosawa, T. Yaita: ISSCC Digest Tech. Papers, (1966) p. 104\n4.22 R. Igarashi, T. Yaita: Proc. AFIPS 1967 FJCC 22, 499 (1967)\n4.23 R.M. Lea: Datafair 73. II, 418 (1973)\n4.24 R.M. Lea: Radio Electron. Eng. 45, 177 (1975)\n4.25 W.F. Bankowski, Jr.: IBM Tech. Discl. Bull. 18,477 (1975)\n4.26 W.F. Bankowski, Jr., K. Simanavicius: IBM Tech. Discl. Bull. 18,475\n(1975)\n4.27 J.L. Mundy, R.E. Joynson: Proc. IEEE Comput. Soc. Conf., (1971) p. 189\n4.28 J.L. Mundy: U.S. Patent No. 3,701,980, Oct. 31 (1972)\n4.29 J.L. Mundy, J.F. Burgess, R.E. Joynson, C. Neugebauer: IEEE JSSC 7,\n364 (1972)\n4.30 R.M. Lea: IEEE JSSC 10, 179 (1975)\n4.31 R.M. Lea: Electr. Eng. 49, 77 (1977)\n4.32 J.R. Burns, J.J. Gibson, A. Harel, K. Hu, R.A. Powlus: RCA Labs. Final\nRept. June 1967 (AD-828 570)\n4.33 M.L. MacKnight: U.S. Patent No. 3,696,174, Sept. 19 (1972)\n4.34 G. Carlstedt, G.P. Petersson, K.o. Jeppson: IEEE JSSC 8, 338 (1973)\n4.35 A. Corneretto: Electron. Des. 2, 40 (1963)\n4.36 C.P. Wang, A.E. Ruehli: IBM Confidential, 65C-061359-MFo03 (1965)\n4.37 R.J. Koerner, S. Nissim: U.S. Patent No. 3,284,775, Nov. 8 (1966)\n4.38 W.K. French: U.S. Patent No. 3,123,706, March 1964\n4.39 W.K. French: U.S. Patent No. 3,131,291, April 1964\n4.40 A.E. Slade, C.R. Smallman: Solid-State Electr. 11, 357 (1960)\n4.41 C.R. Smallman, A.E. Slade, M.L. Cohen: Proc. IRE 48, 1562 (1960)\n4.42 R.W. Ahrons: RCA Rev. 24, 325 (1963)\n4.43 R.W. Ahrons: RCA Rev. 26, 557 (1965)\n4.44 R.W. Ahrons: IEEE Trans. EC-14, 267 (1965)\n4.45 R.W. Ahrons, L.L. Burns: RCA Final Rept., May 1964 (AD-448 504)\n4.46 R.W. Ahrons, L.L. Burns: Comput. Des. 3, 12 (1964)\n4.47 J.L. Anderson: U.S. Patent No. 3,229,255, Jan. 11 (1966)\n4.48 M. Asher: Electron. News 6, 59 (1962)\n4.49 J.D. Barnard, F.A. Behnke, A.B. Lindqvist, R.R. Seeber: Proc. IEEE\n52, 1182 (1964)\n4.50 F.A. Behnke: U.S. Patent No. 3,184,717, May 18 (1965)\n4.51 F.A. Behnke, G.B. Rosenberger: IBM Final Rept. Sept. 1963, (AD-423 492)\n4.52 J.W. Bremer, D.W. Doss, B.T. McKeever: U.S. Patent No. 3,311,898\nMarch 28 (1967)\n4.53 J.W. Bremer, D.W. Doss, B.T. McKeever: U.S. Patent No. 3,312,956\nApril 4 (1969)\n4.54 P.M. Davies: Proc. AFIPS 1962 SJCC 21, 79 (1962)\n4.55 R.J. Ferris: RADC Rept. No. TOR 64-457, Dec. 1964 (AD-61o 131)\n4.56 H. Fleischer, R.I. Roth: U.S. Patent No. 3,221,157, Nov. 30 (1965)\n4.57 General Electric Co.: U.S. Patent No. 3,321,746, May (1967)\n4.58 J. Goldberg, M.W. Green: Stanford Res. Inst. RADC-TR-61-233,\nAug. 1961 (AD-266 169) 340\n4.59 J. Goldberg, M.W. Green: \"Large Files for Information Retrieval Base\non Simultaneous Interrogation of All Items\", in Large-Capacity Memory\nTechniques for Computing Systems, ed. by M.C. Yovits (MacMillan Co.,\nNew York 1962) p. 63\n4.60 K. Goser, H.G. Kadereit: Proc. IEEE 56, 121 (1968)\n4.61 C.C. Green, B. Raphael: Stanford Res. Inst. May 1967 (AD-656789)\n4.62 M.W. Green: Suppl. C. Quart. Rept. 2, AF-30(602)-2142, RADC July 1960\n4.63 M.W. Green: U.S. Patent No. 3,243,785, March 29 (1966)\n4.64 R.S. Green, J. Minker, W.E. Shindle: Auerbach Corp., Management Rept., Vol I, Final Rept. July 1966 (AD-489 660)\n4.65 R.S. Green, J. Minker, W.E. Shindle: Auerbach Corp., Tech. Disc.\nFinal Rept., July 1966 Vol II (AD-489 661)\n4.66 C.H. Heckler, Jr.: In Multiple Instantaneous Response File, p. 195,\ned. by J. Goldberg, RADC-TR-61-233, 1961 (AD-266 169)\n4.67 H.G. Kadereit, K. Goser: Proc. IEEE 56, 121 (1968)\n4.68 A.B. Lindqvist: IBM Tech. Discl. Bull. 7, 1115 (1965)\n4.69 H.T. Mann, J.L. Rogers: Proc. Nat. Aerosp. Electron. Conv. (1962) p. 359\n4.70 W.L. McDermid, R.I. Roth: U.S. Patent No. 3,242,468, March 22 (1966)\n4.71 B.T. McKeever: Proc. AFIPS 1965 FJCC 28, 371 (1965)\n4.72 V.L. Newhouse, R.E. Fruin: Proc. AFIPS 1962 SJCC 21, 89 (1962)\n4.73 V.L. Newhouse, R.E. Fruin: Electronics 35, 31 (1962)\n4.74 J.P. Pritchard: Texas Instruments F.T. Rept. 100~T66, RADC-TR-66-775\n(AD-8ll 983)\n4.75 J.P. Pritchard: Texas Instruments, Final Rept. May 1965 (AD-618 491)\n4.76 J.P. Pritchard: IEEE Spectrum 3, 46 (1966)\n4.77 J.P. Pritchard: IEEE Comput. Group News 2, 25 (1968)\n4.78 J.P. Pritchard, L.D. Wald: Proc. Int. Conf. Nonl inear Magn. (1964)\np. 2-5-1\n4.79 J.P. Pritchard, L.D. Wald: IEEE Trans. MAG-l, 68 (1965)\n4.80 J.A. Rajchman: ONR Rept. ACR-97, Inform. Syst. Summaries, July 1964\n4.81 G. Retiz: IEEE Symp. on Search Memory, May 1964\n4.82 J.L. Rogers: TRW Space Tech. Lab., Quart. Rept. Apr. 1963\n4.83 J.L. Rogers: TRW Space Tech. Lab., Quart. Rept. Aug. 1963\n4.84 J.L. Rogers: ONR Rept. ACR-97, Task No. NR-348-002, RR 003-10-02, 1964\n4.85 J.L. Rogers, A. Wolinsky: TRW Space Tech. Labs., Final Rept.\nNo. NR 3839 (1001), May 1964\n4.86 J.L. Rogers, A. Wolinsky: U.S. Gov. Res. Repts. 39, 166 (1964)\n4.87 H. Rosenberg: U.S. Patent No. 3,235,839, Feb. 15 (1966)\n4.88 G.B. Rosenberger: IBM Data Syst. Div. Final Tech. Rept. 1964 (AD-602 067)\n4.89 R.F. Rosin: Proc. AFIPS 1962 SJCC 21, 203 (1962)\n4.90 P. Schupp, T. Singer: Mitre Corp., Aug. 1963 (AD-416 301)\n4.91 R.R. Seeber, Jr.: Proc. EJCC 18, 179 (1960)\n4.92 R.R. Seeber, Jr.: IBM Data Systems, TR-00,756, Nov. 1960\n4.93 R.R. Seeber, Jr.: Proc. Nat. Conf. ACM 14 (1960)\n4.94 R.R. Seeber, Jr., A.J. Scriver, Jr.: U.S. Patent No. 3,191,155,\nJune 22 (1965)\n4.95 A.E. Slade: Proc. Int. Symp. Theory Switching, (1959) p. 326\n4.96 A.E. Slade: Proc. IRE 50, 81 (1962)\n4.97 Space Technology Laboratories, Inc.: Rept. Proposal 0739.00, July 1961\n4.98 E.D. Van De Rift: In Multiple Instantaneous Response File, p.158, ed. by\nJ. Goldberg, RADC-TR-61-233, 1961 (AD-266 169)\n4.99 1.0. Voitovich: Rept. No. FTD-HT-23-942-68, May 5, 1969 (AD-695 318)\n4.100 C. Yang: \"A Study of Cryotron Associative Memory in Digital Systems\",\nM.Sc. Thesis, Northwestern Univ. (1964)\n4.101 C.C. Yang, J.T. Tou: J. Franklin Inst. 284, 109 (1967)\n4.102 S.S. Yau, C.C. Yang: Proc. Nat. Electronic Conf. 22, 764 (1966)\n4.103 S.S. Yau, C.C. Yang: Northwestern Univ. Tech. Rept., Nov. 1966\n(AD-644 439) 4.104 J. Matisoo: Proc. IEEE 55, 172 (1967)\n4.105 J. Matisoo: IEEE Trans. MAG-5, 848 (1969)\n4.106 W. Anacker: IEEE Trans. MAG-5, 968 (1969)\n4.107 H.H. Zappe: IEEE Trans. MAG-13, 41 (1977)\n4.108 W.Y.Lum, H.W. Chan, T. Van Duzer: IEEE Trans. MAG-13, 48 (1977)\n4.109 P. Wolf: Proc. Int. Conf. SQUID 1976 (de Gruyter, Berlin, New York\n1977) p. 519\n4.110 H.W. Chan, B.T. Ulrich, T. Van Duzer: Proc. Int. Conf. SQUID 1976\n(de Gruyter, Berlin, New York 1977) p. 555\n341\n4.111 T.A. Fulton, L.N. Dunkleberger, R.C. Dynes: Phys. Rev. B 6, 855 (1972)\n4.112 D.J. Herrell: IEEE JSSC 9, 277 (1974)\n4.113 H.H. Zappe: Appl. Phys. Lett. 25, 424 (1974)\n4.114 P. Gueret: IEEE Trans. MAG-ll, 751 (1975)\n4.115 P. Gueret, Th. O. Mohr, P. Wolf: IEEE Trans. MAG-13, 52 (1977)\n4.116 S. Hasuo, T. Imamura, K. Dazai: Proc. Int. Conf. on SQUID 1976\n(de Gruyter, Berlin, New York 1977) p. 541\n4.117 W.F. Bankowski, H.C. Hamel: IBM Tech. Discl. Bull. 16, 3321 (1974)\n4.118 E.I.Il'yashenko, V.F. Rudakov: Assotsiativnye zapominayuzie ustroistva\nna magnitnykh elementakh (Energiya, Moscow 1975)\n4.119 R.T. Hunt, D.L. Snider, J. Surprise, H.N. Boyd: U.S. Gov. Res. Repts.\n39, 188 (1964)\n4.120 H.M. Beisner: IBM Tech. Discl. Bull 8,445 (1965)\n4.121 G.T. Tuttle: U.S. Patent No. 3,300,761, Jan. 24 (1967)\n4.122 Y. Chu: Digital Computer Design Fundamentals (McGraw-Hill, New York 1962)\n4.123 R.C. Corbell: UCLA M.Sc. Thesis (1962)\n4.124 R.R. Seeber, F.B. Hartman: IBM Tech. Discl. Bull. 4, 73 (1962)\n4.125 L.L. Lussier, R.P. Schneider: Electron. Ind. 22, 92 (1963)\n4.126 G.T. Tuttle: Electronics 36, 43 (1963)\n4.127 J.A. Capobianco, J.E. McAteer, R.L. Koppel: Proc. AFIPS 1964 FJCC\n26, 27 (1964)\n4.128 R.J. Koerner, A. Scarborough,: IEEE Symp. Search Memory, May 1964\n4.129 R.J. Koener, A.D. Scarborough: U.S. Patent No. 3,297,995, Jan. 10 (1967)\n4.130 J. McAteer: IEEE Symp. on Search Memory, May 1964\n4.131 A.D. Robbi, R. Ricci: Proc. Int. Conf. Nonlinear Magn. (1964) p. 8-3-1\n4.132 E.S. Lee III: U.S. Patent No. 3,206,735, Sep. 14 (1965)\n4.133 V. Chlouba: Inf. Process. Mach. 13, 113 (1967)\n4.134 J.T. Franks, G.T. Tuttle: U.S. Patent No. 3,300,760, Jan. 24 (1967)\n4.135 R.G. Of engen den , F.N. Berezin: Prib. Tekh. Eksp. No.2, p. 5 (1967)\n4.136 E.I. 11 'yashenko: Autom. Remote Control USSR 34, 1164 (1973)\n4.137 J.A. Rudolph: The Associative Processor, a New Computer Resource,\nGER-14087 (Goodyear Aerospace Corp., Akron, Ohio 1969)\n4.138 J.A. Rudolph, L.C. Fulmer, W.C. Meilander: Electronics 43,96 (1970)\n4.139 J.A. Rudolph, L.C. Fulmer, W.C. Meilander: Electronics 44, 91 (1971)\n4.140 R.H. Fuller, J.D. Tu, R.M. Bird: Proc. Nat. Aerosp. Electron. Conf.\n(1965) p. 1\n4.141 R.J. Koerner: U.S. Patent No. 3,257,650, June 21 (1966)\n4.142 W.F. Chow: IEEE Trans EC-16, 642 (1967)\n4.143 W.F. Chow, L.M. Spandorfer: Proc. AFIPS 1967 SJCC 28 (1967)\n4.144 M. Bialer, J. Garrett, W.C. Meilander: Symp. Parallel Processor\nSystems, Tech. & Appl., June 1969\n4.145 J.P. Mc Callister, C.F. Chong: Proc. AFIPS 1966 FJCC 29 (1966)\n4.146 A. Corneretto: Electron. Des. 10, 8 (1962)\n4.147 C.A. Rowland, W.O. Berge: Proc. AFIPS 1963 FJCC 24, 59 (1963)\n4.148 J.I. Raffel, T.S. Crowther: IEEE Trans. EC-13, 611 (1964)\n4.149 M. Naiman: \"Content-Addressed Memory Using Magneto-Resistive Readout\nof Magnetic Thin Films\", IEEE Intermag. 1965\n4.150 C.P. Wang: J. Appl. Phys. 39, 1220 (1968)\n4.151 C.P. Wang: U.S. Patent No. 3,466,631, Sept. 9 (1969) 342\n4.152\n4.153\n4.154\n4.155\n4.156\n4.157\n4.158\n4.159\n4.160\n4.161\n4.162\n4.163a\n4.163b\n4.164\n4.165\n4.166\n4.167\n4.168\n4.169\n4.170\n4.171\n4.172\n4.173\n4.174\n4.175\n4.176\n4.177\n4.178\n4.179\n4.180a\n4. 180b\n4.181\n4.182\n4.183\n4.184\n4.185\n4.186\n4.187\n4.188\n4.189\nC.P. Wang: U.S. Patent No. 3,466,632, Sept. 9 (1969)\nP.A. Lord, M.P. Marcus: U.S. Patent No. 3,426,335, Feb. (1969)\nJ. Kiseda, H. Petersen, W. Seelbach, M. Teig: IBM J. Res. Dev. 5,\n106 (1961)\nJ. R. Ki seda: \"A 128-Word, 36-Bit Magneti c Associati ve Memory\". IBM\nconfidential, NC-358, March 1964\nW.I. McDermid, J.E. Petersen: IBM J. Res. Dev. 5, 59 (1961)\nR.G. Gall: Proc. AFIPS 1964 FJCC, p. 159 (1964)\nA. Api ce 11 a, J. Franks: \"BILOC - A Hi gh Speed NDRO One-Core-Per-Bit\nAssoci ati ve El ement\", IEEE Intermag. 1965\nTse-Yun Feng: Proc. AFIPS 1968 SJCC 33, 275 (1968)\nC.F. Pulvari, M. Szabo, M.J. Walsh, A. De La Paz, W.B. Penzes:\nRADC-TR-68-105, April 1968 (AD-668 475)\nC.F. Pulvari: IEEE Trans. ED-16, 580 (1969)\nR.F. Herlein, A.V. Thompson: Proc. Int. Solid State Circuits Conf.\n(1969) p. 42\nD. Toombs: IEEE Spectrum 15, 22 (1978)\nD.F. Barbe: Charge-Coupled Devices, Topics in Applied Physics, Vol. 38 (Springer, Berlin, Heidelberg, New York 1980)\nC. Kooy, U. Enz: Philips Res. Rep. 15, 7 (1960)\nA.H. Bobeck: Bell Syst. Tech. J. 46, 1901 (1967)\nH. Chang: Magnetic-Bubble Memory Technology (Dekker, New York, Basel\n1978)\nA.M. Bobeck, H.E.D. Scovil, W. Shockley: U.S. Patent 3,541,522\n(1967, Pat. 1970)\nH. Murakami: U.S. Patent 3,760,390 (1972, Pat. 1973)\nS.Y. Lee, H. Chang: IEEE Compo Soc. Conf. COMPCON 75 1, 91 (1975)\nW. Kluge: Proc. IEEE 120, 1308 (1973)\nI.G. Avaeva, E.I. Il'yashenko, V.G. Kleparskii, Yu.L. Kopylov,\nV.B. Kravchenko, A.T. Sobolev: Mikroelektronika (USSR) 5, 188 (1976)\nI.G. Avaeva, E.I. Il'yashenko, Yu.L. Kopylov, V.B. Kravchenko,\nF.V Lisovskii, S.N. Matveev, A.T. Sobolev, V.I. Shapovalov: Mikroelektronika (USSR) 5, 500 (1976)\nR. Allan: IEEE Spectrum 12, 49 (1975)\nD.M. Lee, R.A. Naden: NAECON '76 Record, p. 724 (1976)\nE.I. Il'yashenko, I.G. Avaeva, V.B. Kravchenko, S.N. Matveev:\nPhys. Status Solidi A: 36, Kl (1976)\nE.I. 11 'yashenko, V.G. Kleparskii, S.E. Yurchenko: Phys. Status\nSolidi A: 28, K153 (1975)\nE.I. 11 'yashenko, S.N. Matveev, N.I. Karmatzky: IEEE Trans.\nMAG-12, 663 (1976)\nV.G. Kleparsky, E.I. Il'yashenko, S.N. Matveev: IEEE Trans.\nMAG-12, 700 (1976)\nE.I. 11 'yashenko, S.N. Matveev: Zh. Tekh. Fiz. 3, 137 (1977)\nR.A. Naden: U.S. Pat. Appl. 809 729 (AD-D004 142/6S) (1977)\nA.H. Eschenfelder: Magnetic Bubble Technology, Springer Series in Solid\nState Sciences, Vol.14 (Springer, Berlin, Heidelberg, New York 1980)\nD.O. Smith, K.J. Harte: IEEE Trans. EC-15, 123 (1966)\nD.O. Smith, K.J. Harte: IEEE Trans. EC-16, 372 (1967)\nK.Y. Ahn, Y.S. Lin: IBM Tech. Discl. Bull. 15, 2017 (1972)\nR.J. Collier: IEEE Spectrum 3, 67 (1966)\nG.W. Stroke: Appl. Phys. Letters 6, 201 (1965)\nG.W. Stroke: An Introduction to Coherent Optics and Holography\n(Academic Press, New York 1966)\nW. Kulcke, K. Kosanke, E. Max, M.A. Habegger, T.J. Harris, H. Fleischer:\nProc. IEEE 54, 1419 (196b)\nH. Kogelnik: Microwaves 6, 68 (1967)\nD.R. Bosomworth, H.J. Gerritsen: Appl. Opt. 7, 95 (1968) 4.190\n4.191\n4.192\n4.193\n4.194\n4.195\n4.196\n4.197\n4.198\n4.199\n4.200\n4.201\n4.202\n4.203\n4.204\n4.205\n4.206\n4.207\n4.208\n4.209\n4.210\n4.211\n4.212\n4.213\n5.1\n5.2\n5.3\n5.4\n5.5\n5.6\n5.7\n5.8\n5.9\n5.10\n5.11\n5.12\n5.13\n5.14\n5.15\n5.16\n5.17\n5.18\n5.19\n5.20\nM. Sakaguchi, N. Nishida: U.S. Patent 3,704,929 (1970, Pat. 1972)\nG.R. Knight: Appl. Opt. 13,904 (1974)\nP.P. Sorokin: IBM J. Res. Dev. 8, 182 (1964)\nM. Sakaguchi, N. Nishida, T. Nemoto: IEEE Trans. C-19, 1174 (1970)\nG.R. Knight: Appl. Opt. 14, 1088 (1975)\nJ.T. LaMacchia, D.L. White: Appl. Opt. 7,91 (1968)\nJ. Thire: L'Onde Electr. 52, 452 (1972)\n343\nV. Malina, M. Chomat: Slaboproudy (Czechoslovakia) 37, 572 (1976)\nN. Minnaja, M. Nobile: Rend. Riunione Annu. Assoc. Elettrotec. Ital.\n49, B3/1 (1974) .\nN. Minnaja, M. Nobile: Elettrotecnica 61, 674 (1974)\nN. ~1innaja: U.S. Patent No. 3,887,906, June 3 (1975)\nD. Garelli: Elektrotechnik 57, 30 (1975)\nN. Nishida, M. Sakaguchi, F. Saito: Appl. Opt. 12, 1663 (1973)\nG.N. Aleksakov, Yu.A. Bykhovskii, A.I. Larkin, A.A. Markilov:\nPrib. Tekh. Eksp. (USSR) 17, 203 (1974)\nM.T. Fatehi, S.A. Collins Jr.: Int. Opt. Comput. Conf. (Digest of\nPapers), Apr. 23-25, 1975, p. 105\nV.K. Bykhovsky, A.K. Glotov, A.E. Krasnov: Int. Opt. Comput. Conf.,\nCapri, Italy, Aug. 28 - Sep. 2, 1976\nE. MUhlenfeld: Int. Opt. Comput. Conf., Capri, Italy, Aug. 28 - Sep. 2,\n1976 (Digest of Papers) p. 111\nE. MUhlenfeld: Complete manuscript of [4.206]\nE. MUhlenfeld, W. Kistner, A. Keller: Elektron. Rechenanlagen 18,\n224 (1976)\nH. Akahori: Res. Electrotech. Lab. (Japan), No. 771, p. 1, Oct. 1977\nI.V. Prangishvili, A.K. Glotov, A.E. Krasnov, V.K. Bykhovskii: Proc.\n2nd USA-USSR Seminar Opt. Inf. Processing Novosibirsk, July 1976\n(Plenum Press, New York 1977)\nV.N. Morozov: SOY. J. Quantum Electron. (USA) 8, 1 (1978)\nA.A. Vasiliev, I.N. Kompanet, S.P. Kotova, V.N. Morozov: Kvantovaya Elektron. 5, 1298 (1978)\nA.A. Vasiliev: Proc. 1st Eur. Conf. Opt. Syst. and Appl. 4-6 Apr. 1978, Brighton, England, Vol. 6 (1978)\nC.J. Conti: IEEE Comput. Group News 2, 9 (1969)\nR.M. Meade: Electronics 45, 58 (1972)\nR.M. Meade: IFIPS 1970 FJCC, p. 33\nL.A. Belady: IBM Syst. J. 5, 78 (1966) R.~1. Meade: Comput. Des. 10, 87 (1971)\nI.S. Reed: Computer 5, 47 (1972)\nD.M. Nessett: Aust. Comput. J. 7, 33 (1975)\nD.P. Agrawal: Proc. 1977 Nat. Comp. Conf., 13-16 Jun. (1977)\nD.P. Agrawal, R.J. Zingg, A.V. Pohm: Proc. 14th IEEE Comput. Soc.\nInt. Conf. (IEEE, New York 1977) p. 74\nJ.D. Jones, D.M. Junod, R.L. Partridge, B.L. Shawley: IBM Tech. Discl.\nBull. 19, 594 (1976)\nJ.D. Jones, D.M. Junod: IBM Tech. Discl. Bull. 20, 295 (1977)\nT. Kilburn, D. Edwards, N. Lanigan, F.H. Summer: IEEE Trans. EC-ll,\n233 (1962)\nS.G. Campbell: Proc. AFIPS 1963 FJCC 24, 473 (1963)\nM.F. Wolff: Electronics 36, 35 (1963)\nG.G. Scarrot: Proc. IFIP Congo 1965, 1, 137\nL.C. Hobbs: IEEE Trans. EC-15, 534 (1966)\nA.B. Lindqvist, R.R. Seeber, L.W. Comeau: Proc. IEEE 54, 1774 (1966)\nW. Anacker, C.P. Wang: IEEE Trans. EC-16, 764 (1967)\nE.J. Joseph: Comput. Des. 8, 165 (1969)\nR.L. Mattson, J. Gecsei, D.R. Slutz, I.L. Traiger: IBM Syst. J. 2,\n78 (1970) 344\n5.21 H. Katzan, Jr.: Proc. 1971 SJCC, p. 325\n5.22 J.G. Williams: Commun. ACM 14, 172 (1971)\n5.23 A. Bensoussan, C.T. Clingern, R.C. Daley: Commun. ACM 15, 308 (1972)\n5.24 R.P. Parmelee, T.I. Peterson, C.C. Tillman, D.J. Hatfield: IBM Syst.\nJ. 11, 99 (1972)\n5.25 R.E. Brundage, A.P. Batson: Proc. 2nd Annu. Symp. Comput. Archit.,\nJan. 20-22, 1975 (IEEE, New York 1975) p. 85\n5.26 R.F. Brundage, A.P. Batson: Rev. Fr. Automat. Inf. Rech. Oper. 10, 47\n(1976)\n5.27 S.Ya. Berkovich, Yu.Ya. Kochin, Yu.N. Khrebtov: Program. and Comput.\nSoftware (USA) 2, 455 (1976)\n5.28 C. Schuenemann: IBM Tech. Discl. Bull. 21,663 (1978)\n5.29 T.D. Chase, R.M. Glorioso: Proc. 25th Anniv. Conf. 1, 6 (1972)\n5.30 M.V. Wilkes: IEEE Trans. EC-14, 270 (1965)\n5.31 M.V. Wilkes: IEEE Trans. C-20, 674 (1971)\n5.32 D.H. Gibson, W.L. Shevel: Electronics 21, 105 (1969)\n5.33 F.F. Lee: IEEE Trans. C-1B, 1062 (1969)\n5.34 D.C. Gunderson: Computer 3, 7 (1970)\n5.35 H.S. Stone: IEEE Trans. C-19, 73 (1970)\n5.36 J.H. Kroeger, R.M. Meade: Proc. 1971 Comput. Des. Conf., Vol. 1 (1971)\np. 252\n5.37 H. Barmasian, A.L. DeCegama: Digest of Papers 6th Annu. IEEE Comput. Soc. Int. Conf., 1972, p. 107\n5.38 K.R. Kaplan, R.O. Winder: Computer 6, 30 (1973)\n5.39 J. Niederreichholz: Elektron. Rechenanlagen 1B, 122 (1976)\n5.40 B.D. Ackland, D.A. Pucknell: Electron. Lett. 11,588 (1975)\n5.41 B.D. Ackland, D.A. Pucknell: N76/5 For. Meet.: Comput. in Eng., Perth\nSept. 6-7, 1976, p. 120\n5.42 A.J. Smith: Proc. 2nd Int. Conf. Software Eng. (IEEE, New York 1976)\np. 286\n5.43 A.J. Smith: IEEE Trans. SE-4, 121 (1978)\n5.44 H. Iizuka, T. Terui: J. Inf. Proc. Soc. Jpn. 14, 669 (1973)\n5.45 M. Bennett, P. Berard, C. Boksenbaum, M. Veran: Rev. Fr. Automat. Inf.\nRech. Oper., 10 May,1976, p. 1\n5.46 C.K. Chow: IBM Tech. Discl. Bull. 17, 3163 (1975)\n5.47 C.K. Chow: IBM Tech. Discl. Bull. 1B, 1643 (1975)\n5.48 C.K. Chow: IEEE Trans. C-25, 157 (1976)\n5.49 C.G. Bell, D.P. Casasent: Comput. Des. No. II, Nov. 1971, p. 83\n5.50 J. Bell, D.P. Casasent, C.G. Bell: IEEE Trans. C-23, 346 (1974)\n5.51 R. Monroe: Digest of Papers 10th IEEE Compo Soc. Int. Conf. 3, 1034\n(1975)\n5.52 W.O. Strecker: Proc. 3rd Annu. Symp. Comput. Archit. (IEEE New York\n1976) p. 155\n5.53 A. Weinberger: U.K. Patent No. 1 280 753, July 5 (1972)\n5.54 O.K. Chia: IBM Tech. Discl. Bull. 17, 3361 (1975)\n5.55 A.M. Weiner: IBM Tech. Discl. Bull. 19, 4697 (1977)\n5.56 C.J. Conti, D.H. Gibson, S.H. Pitkowsky: IBM Syst. J. 7, 2 (1968)\n5.57 J.S. Liptay: IBM Syst. J. 7, 15 (1968)\n5.58 R.A. McLaughlin: Datamation, Sept. 1972, p. 58\n5.59 Compo Des. 15. 26 (1976)\n5.60 J.F. Mastranadi: IBM Tech. Discl. Bull. 14, 3760 (1972)\n5.61 M.J. Haims: IBM Tech. Discl. Bull. 18. 278 (1975)\n5.62 H. Gelernter: Proc. Symp. on Large Capacity Memory Tech .• May 1961\n5.63 H.E. Petersen: Proc. IFIPS Cong., Aug. 1962\n5.64 H. Hellerman, IBM Watson Res. Center Rept. No. RC-I095, Oct. 1963\n5.65 F.J. Hilbing: Ph. D. Dissert., Stanford Univ. (1968)\n5.66 J.B. Rothnie. T. Lozano: Commun. ACM 17, 63 (1974)\n5.67 G.M. Furxhi: Riv. Inf. (Italy) 7, 15 (1977) 5.68\n5.69\n5.70\n5.71\n5.72\n5.73\n5.74\n5.75\n5.76\n5.77\n5.78\n5.79\n5.80\n5.81\n5.82\n5.83\n5.84\n5.85\n5.86\n5.87\n5.88\n5.89\n5.90\n5.91\n5.92\n5.93\n5.94\n5.95\n5.96\n5.97\n5.98\n5.99\n5.100\n5.101\n5.102\n5.103\n5.104\n5.105\n5.106\n5.107\n5.108\n5.109\nP.B. Berra: Proc. COMPSAC 1978 (IEEE, New York 1978) p. 698\nA.J. Symonds: IBM Syst. J. 7, 229 (1968)\nJ. Fotheringham: Commun. ACM 4, 435 (1961)\n345\nD. Aspinall, D.J. Kinniment, D.B.G. Edwards: IFIP Edinburgh, Aug. 1968,\np. 081\nH.R. Holt, J.A. Timmons, D.C. Gunderson: Honeywell Inc., Final Tech.\nRept. 12099-FR1, Sept. 1968\nM.H.J. Baylis, D.G. Fletcher, D.J. Howarth: Inform. Froc. 68\n(North-Holland, Amsterdam 1969) p. 831\nY. Chu: Computer Organization and Microprogramming (Prentice-Hall, Englewood Cliffs, N.J. 1972)\nR. Moulder: Proc. AFIPS Nat. Conf. Comput. Composition and Expo. 42,\n171 (1973)\nW.B. Riley: Electronics 45, 91 (1972)\nJ.L. Gertz: Infotech. State-of-the-Art Rept. (Maidenhead, England 1976) p. 273\nK. Koch: IBM Tech. Discl. Bull. 15, 3088 (1973)\nJ.R. Carlberg: Taylor Naval Ship Res. and Dev. Center, Rept. No.\nDTNSRDC-77-0083, Aug. 1977\nM. Takesue: Inform. Process. Soc. Jpn. 19, 158 (1978)\nC.V.W. Armstrong: Proc. 2nd Annu. Symp. Comput. Archit. (IEEE, New\nYork 1975) p. 34\nC.E. Shannon: Bell Syst. Tech. J. 28, 59 (1949)\nT.F. Tabloski, F.J. Mowle: IEEE Trans. C-25, 684 (1976)\nW.E. Donath: IBM J. Res. Dev. 18, 401 (1974)\nB.A. Holum: IBM Confidential, SRI Term Paper, No. 11-31, April 1964\nF.T. Baker, W.E. Triest, C.H. Forbes, N. Jacobs, J. Schenken: IBM,\nFinal Rept. May 1966, AF-30(602)-3573\nD.C. Gunderson, J.P. Francis, W.L. Heimerdinger: Honeywell Inc.,\nRept. No. 12029, Dec. 1966 (RADC TR-66-573)\nD.C. Gunderson, W.L. Heimerdinger, J.P. Francis: \"A Multiprocessor\nwith Associative Control\", in Frospects for Simulation and Simulators\nof Dynamic Systems (Spartan Books, New York 1967) p. 183\nR. Gonzales, D.C. Gunderson, J.A. Timmons: Honeywell Inc., Final\nRept. Nov. 1967 (AD-662 361)\nR.P. Bair: Moore School of Electr. Eng., May 1968 (AD-674 199)\nL.D. Wald, G.A. Anderson: Final Rept. NAS 12-2087, Sept. 1971\nF. Tsui: IBM Tech. Discl. Bull. 15,2342 (1972)\nL. Hellerman, G.E. Hoernes: IEEE Trans. C-17, 1144 (1968)\nI.N. Hooton: In Automatic Acquisition and Reduction of Nuclear Data\n(Ges. FUr Kernforschung G.m.b.H., Karlsruhe 1964) p. 338\nLN. Hooton: In Ref. 5.94, p. 349\nH. Meyer, W. Stuber: In Ref. 5.94, p. 357\nE. Blanca, A. Carriere: CEA-R-3394, Dec. 1967\nM.D. Johnson, D.C. Gunderson: Proc. 1970 Int. Telemetry Conf.,\nApril 1970, p. 109\nL. Rettelbusch, H. Pfahlbusch: Nachrichtentech. Elektron. 24, 340\n(1974)\nT.L. Saxton, C.-C. Huang: IEEE Trans. C-26, 170 (1977)\nR.R. Seeber, Jr.: Commun. ACM 4, 301 (1961)\nLS. Gershuny, O.L. Lamb: IBt~ Tech. Discl. Bull. 15, 1109 (1972)\nB.A. Crane: IEEE Trans. C-17, 691 (1968)\nB.H. Scheff: Electron. Prog. 10, 31 (1966)\nC. Peters: NTIS AD-824 213\nS.N. Porter: J. ACM 13, 369 (1966)\nR.C. Minnick: IEEE Trans. EC-13, 685 (1964)\nC.C. Yang, S.S. Yau: IEEE Trans. EC-15, 522 (1966)\nS.S. Yau, M. Orsic: IEEE Trans. C-19, 259 (1970) 346\n5.110 S.S. Yau, C.K. Tang: IEEE Trans. C-19, 141 (1970)\n5.111 C. Barre: Electron. Appl. Ind. (France) 250, 21 (1978)\n6.1 C.Y. Lee, M.C. Paull: Proc. IEEE 51, 924 (1963)\n6.2 C. Lee, M. Paull: Proc. IEEE 52,312 (1964)\n6.3 C. Y. Lee: \"Content-Addressable and Distributed Logic Memories\", in\nApplied Automata Theory, ed. by J.T. Tou (Academic Press, New York\n1968)\n6.4 E.S. Lee: Proc. AFIPS 1963 SJCC, p. 381 (1963)\n6.5 R.S. Gaines, C.Y. Lee: IEEE Trans. EC-14, 72 (1965)\n6.6 B.A. Crane, J.A. Githens: IEEE Trans. EC-14, 186 (1965)\n6.7 G. Nemeth: Helsinki U.Tech., Dept. Tech. Phys. Report TKK-F-A347\n( 1978)\n6.8 B.A. Crane, R.R. Laane: Proc. AFIPS 1967 SJCC, p. 517 (1967)\n6.9 R.P. Edwards: Proc. IEEE 52, 83 (1964)\n6.10 E.S. Spiegelthal: Proc. IEEE 52, 74 (1964)\n6.11 A. Tremblay: Cybernetics XIX, 105 (1976)\n6.12 J.N. Sturman: IEEE Trans. C-l?, 2 (1968)\n6.13 J.N. Sturman: IEEE Trans. C-l?, 10 (1968) 6.14 J.E. Smathers: Ph. D. Dissert., Oregon State Univ. (1969)\n6.15 W.H. Kautz, K.N. Levitt, A. Waksman: IEEE Trans. C-l?, 443 (1968)\n6.16 W.H. Kautz: J. ACM 18, 19 (1971)\n6.17 W.H. Kautz, M.C. Pease III: AD 763 710 (1971)\n6.18 J. Hood, M. Mark, J. Cotton: Proc. 1976 Int. Conf. Parallel Processing, Aug. 24-27, 1976, p. 168\n6.19 R. Trepp: RADC-TR-66-182, June 1966\n6.20 C.A. Finnila, H.H. Love, Jr.: IEEE Trans. C-26, 112 (1977)\n6.21 S. Ya. Berkovich, Ya.Ya. Kochin, G.M. Lapir: Autom. Remote Control\n35, 1342 (1974)\n6.22 D.A. Savitt, H.H. Love, R.E. Troop: AD 488 538 (1966)\n6.23 D.A. Savitt, H.H. Love, R.E. Troop, R.A. Rutman: Association Storing\nProcessor Interpretive Program - Program Logic Manual. Final Report,\nHughes Aircraft Co., FR-11-558 (1968)\n6.24 H.H. Love, D.A. Savitt: RADC-TR-65-32 (1965)\n6.25 H.H. Love, D.A. Savitt: In Associative Information Techniques, ed.\nby E.L. Jacks (American Elsevier, New York 1971) p. 147\n6.26 H.H. Love, R.A. Rutman: Hughes Aircraft, FR-68-11-1179, Dec. 1968\n6.27 H.H. Love: Hughes Aircraft, FR-69-11-487, Jun. 1969\n6.28 R.A. Rutman: Hughes Aircraft, FR-69-11-208, Feb. 1969\n6.29 J.H. Holland: 1959 EJCC, p. 108\n6.30 G.J. Lipovski: Proc. AFIPS 1970 SJCC, p. 385 (1970)\n6.31 G.H. Schmitz: Final Rep. Contr. No. DAH 60-72-C0050 (1972)\n6.32 Proc. 1972 Sagamore Compo Conf. (IEEE, New York 1972)\n6.33 W.S. Litzler: 1973 Swieeeco Record of Technical Papers, p. 482\n6.34 E.C. Stanke II: RADC-TR-77-366 (1978) 6.35 J.A. Githens: \"An Associative, Highly-Parallel Computer for Radar\nData Processing\", in Pa:rallel Processor Systems, Technologies, and\nApplications, ed. by L.C. Hobbs, D.J. Theis, J. Trimble, H. Titus,\nI. Highberg (Spartan Books, New York 1970)\n6.36 R.O. Berg, M.D. Johnson: Proc. IEEE 1970 Int. Compo Group. Conf.,\nWashington, p. 336\n6.37 J.A. Githens: Proc. NAECON 1970, p. 290\n6.38 J.A. Githens: Proc. IEEE 1972 Int. Compo Soc. Conf., p. 57\n6.39 R.O. Berg, H.G. Schmitz, S.J. Nuspl: Proc. NAECON 1972, p. 312\n6.40 J.A. Cornell: Proc. WESCON 1972, p. 1/3-1\n6.41 J.A. Cornell: Proc. COMPCON 1972, p. 69\n6.42 K.E. Batcher: WESCON T.ech. Papers 16, 1 (1972)\n6.43 J.A. Rudolph: Proc. AFIPS 1972 FJCC, p. 229 (1972) 6.44 K.E. Batcher: Proc. 1974 Nat. Compo Conf., p. 405\n6.45 E.W. Davis: Proc. 1974 Nat. Compo Conf., p. 17\n6.46 Goodyear Aerospace Corp.: Doc. 8284 C (1974)\n6.47 Goodyear Aerospace Corp.: Doc. GER-15636B (1974)\n6.48 Goodyear Aerospace Corp.: Doc. GER-15637B (1974) 6.49 Goodyear Aerospace Corp.: Doc. GER-15643A (1974)\n6.50 Goodyear Aerospace Corp.: Doc. GER-15644A (1974)\n6.51 Goodyear Aerospace Corp.: Doc. GER-16139 (1974)\n6.52 Goodyear Aerospace Corp.: Doc. AP-112286 (1965)\n347\n6.53 D. Brotherton, S. Domchick: Goodyear Aerospace Corp., Doc. GER-12318\n(1966) 6.54 L.C. Fulmer, W.C. Meilander: Proc. 1970 IEEE Int. Compo Group. Conf.,\np. 325\n6.55 W.C. Meilander, J. Garrett, M. Bialer: \"A Mission oriented associative\nprocessor using plated wire\", in ParaUel Proeessor Systems, Teelmol\u0002ogies, and Applieations, ed. by L.C. Hobbs, D.J. Theis, J. Trimble,\nH. Titus, I. Highberg (Spartan Books, New York 1970) p. 153\n6.56 W. Shooman: Proc. 1960 EJCC, p. 111\n6.57 W. Shooman: U.S. Patent 3,277,449 (1966)\n6.58 W. Shooman: \"Orthogonal Processing\", in ParaUel Proeessor Systems,\nTeehnologies, and Applieations, ed. by L.C. Hobbs,D.J. Theis, J. Trimble,\nH. Titus, I. Highberg (Spartan Books, New York 1970)\n6.59 L.C. Highbie: COMPCON '72, p. 287\n6.60 J.C. Murtha, R.L. Beadles: ONR Rep. No. 4755 (1964)\n6.61 J.C. Murtha: \"Highly Parallel Information Processing Systems\", in\nAdvanees in Computers, Vol. 7 (Academic Press, New York 1966) p. 1\n6.62 M.J. Flynn: Proc. IEEE 54, 1901 (1966)\n6.63 G.L. Hollander: Proc. AFIPS 1967 SJCC, p. 463 (1967)\n6.64 L.C. Hobbs, D.J. Theis, J. Trimble, H. Titus, I. Highberg (eds.):\nParallel Proeessor Systems, Teehnologies, and Applieations (Spartan Books, New York 1970)\n6.65 L.C. Higbie: Computer 6, 48 (1973) 6.66 J.E. Shore: 1972 IEEE Int. Conv. Digest, p. 358\n6.67 J.E. Shore: Comput. Electron. Eng. 1, 95 (1973)\n6.68 C.C. Foster: Content Addressable Parallel Proeessors (Van Nostrand,\nNew York 1976)\n6.69 K.J. Thurber: Large Seale Computer Arehiteeture (Hayden, Rochelle\nPark, N.J. 1976)\n6.70 L.C. Roberts: IEEE Spectrum 11, 46 (1974)\n6.71 B.H. McCormick: IEEE Trans. C-12, 791 (1963)\n6.72 D.L. Slotnick: Proc. AFIPS 1967 SJCC, p. 477 (1967)\n6.73 R.M. Barnes, R.M. Brown, M. Kato, D.J. Kuck, D.L. Slotnick, R.A. Stokes:\nIEEE Trans. C-17, 746 (1968)\n6.74 P. Alsberg, J. Gaffney, C. Grossman, T. Mason: Illinois Univ.\nIlliac-IV-212, March 1969\n6.75 R.L. Davis: IEEE Trans. C-18, 800 (1969)\n6.76 B.H. McCormick, J.L. Divilbiss: Rept. No. 4031, Digital Compo Lab.,\nUniv. of Illinois, 1969\n6.77 Burroughs Corp.: ILLIAC IV Systems Charaeteristies and Programming\nManual, Contract Report AF 30 (602) 4144 (1973)\n6.78 E.J. Jacks (ed.): Assoeiative Information Teelmiques (American Elsevier,\nNew York 1971)\n6.79 Proc. 1973 Sagamore Comput. Conf. (IEEE, New York 1973)\n6.80 Proc. 1974 Sa9amore Comput. Conf. (IEEE, New York 1974)\n6.81 Proc. 1975 Sagamore Comput. Conf. (IEEE, New York 1975)\n6.82 M. Feilmeyer (Ed.): Parallel Computers - Parallel Mathematies. Proc.\nIMACS (AICA)-61 Symposium March 14-16, 1977, TU Munich (North-Holland, Amsterdam 1977) 348\n6.83\n6.84\n6.85\n6.86\n6.87\n6.88\n6.89\n6.90\n6.91\n6.92\n6.93\n6.94\n6.95\n6.96\n6.97\n6.98\n6.99\n6.100\n6.101\n6.102\n6.103\n6.104\n6.105\n6.106\n6.107\n6.108\n6.109\n6.110\n6.111\n6.112\n6.113\n6.114\nProc. 1976 Int. Conf. on Parallel Processing, Aug. 24-27, 1976\n(IEEE, New York 1976)\nProc. 1977 Int. Conf. on Parallel Processing, Aug. 23-26, 1977\n(IEEE, New York 1977)\nControl Eng. 9, 22 (1962)\nR.H. Fuller: General Precision-Librascope Inc., Interim Rept., AD-608 427, October 1964\nR.H. Fuller: Comput. Des. 6, 43 (1967)\nR.H. Fuller: Proc. AFIPS 1967 SJCC, p. 471\nR.H. Fuller: General Precision, ONR/RADC Seminar on Assoc. Proc. 1967\nR.H. Fuller, R.M. Bird, J.N. Medick: \"Associative Processor Study\",\nLibrascope Div. General Precision, Oct. 1964\nR.H. Fuller, R.~1. Bird, R.M. Worthy: RADC-TR-65 210, AD-621 516,\nAugust 1965\nWestinghouse Defense and Space Center, Final Rept. June 1964,\nAD-602 693\nJ.A. Feldman: M.I.T. Lincoln Lab. Tech. Note 1965-13, April 1965, AD-614 634\nGeneral Precision Inc.: \"Associative Processing Techniques\" (Librascope Group, 1965)\nD.L. Reich: \"Associative Memories and Information Retrieval\",\nin Some Problems in Information Science, ed. by M. Kochen\n(Scarecrow Press, New York 1965)\nJ.A. Dugan, R.S. Green, J. Minker, W.E. Shindle: Proc. ACM 21st\nNat. Conf. 1966, p. 347\nK.E. Knight: Datamation 12, 40 (1966)\nM.A. Knapp: \"RADC Programs in Associative Processing\", ONR/RADC\nSeminar on Assoc. Proc., May 1967\nH.I. Jauvits: Interim Rept., Lab. For Electronics Inc., FFB, 1968,\nNASA-CR-86076\nJ.A. Rudolph: Proc. IEEE Region 6 Conf., Apr. 1969, p. 179\nM.H. Cannell, A.J. Nickelson, ~1.F. Owens, K.W. Wadman, M.L. Urban:\nMitre Corp., Repts. Nos. MTR-1735-Rev-1, MTR-863, AD-879 281\nDec. 1970\nL.C. Hobbs, D.J. Theis: \"Survey of Parallel Processor Approaches and\nTechniques\", Symp. on Parallel Proc. Systems Technologies and\nApplications, Monterey 1969 (Papers ed. by L.C. Hobbs et al. 1970)\nJ.C. Murtha: NAECON '70 Records, May 1970, p. 298\nW.C. Meilander, R.G. Gall: \"Evaluation of the Goodyear associative\nprocessor in an operational ATC environment\", IEEE Compo Soc. Conf.,\nBoston, Mass., Sep. 1971\nM. Minsky, S. Papert: \"On Some Associative, Parallel and Analog\nComputations\", in Associative Information Techniques, ed. by E.J.\nJacks (American Elsevier, New York 1971)\nK.J. Thurber, R.O. Berg: Comput. Des. 10, 103 (1971)\nB. Parhami: \"Design Techniques for Associative Memories and Processors\",\nUCLA, Comput. Sci., Rept. No. PB-220 714 (1973)\nK.J. Thurber, P.C. Patton: IEEE Trans. C-22, 1140 (1973)\nR.M. Lea: Computer 8, 25 (1975)\nL.C. Higbie: Comput. Electr. Eng. 2, 397 (1975)\nL.C. Higbie: Comput. Des. 15, 75 (1976)\nB.W. Prentice, R. Katz, R. Komadja, H. Lee: Boeing Compo Services Inc.,\nSeattle, Wash. Jan. 1975, RADC-TR-74-326, AD-A005 308\nK.J. Thurber, L.D. Wald: Comput. Surv. 7, 215 (1975)\nD. Lewin: \"Introduction to Associative Processors\", Proc.Conf. Compo\nArchit., St. Raphael, France, 12-24 Sept. 1976, ed by G.G. Boylaye, D.W. Lewin 349\n6.115 M.W. Summers: Rome Air Devel. Cent. Rept. RADC-TR-75-318, Jan. 1976,\nAD-A021 232\n6.116 FToc. IEEE 1977 Int. Conf. on Parallel FTocessing, Aug. 23-26, 1977\n(IEEE, New York 1977)\n6.117 Infotech. Int.: Future Systems, State of the Art Rept. (Maidenhead,\nEngl and, 1977)\n6.118 S.S. Yau, H.S. Fung: Comput. Surv. 9, 3 (1977)\n6.119 N.J. Zimmerman, H.J. Sips: Informatie (Netherlands) 20, 3 (1978)\n6.120 D.L. Slotnick, W.C. Borck, R.C. McReynolds: Proc. AFIPS 1962 FJCC 22,\n97 (1962)\n6.121 Westinghouse Defence and Space Center: \"Parallel Network Computer\n(SOLOmN) Applications Analyses\", August 1964, AD-606 578\n6.122 F.W. Weingarten, P.T. Rux, J.A. Boles: \"On an Associative ~lemory for\nNebula Computer\", Dept. of Math., Oregon State Univ., In-House Doc.,\n1964\n6.123 J.A. Boles: \"The Logical Design of the Nebula Computer\"; Ph. D. Thesis,\nOregon State Univ. (1968) AD-673 990\n6.124 IBM: \"Project Lightning\", AD-250 678 (1960)\n6.125 IBM: \"Project Lightning\", U.S. Gov. Res. Repts. 36, 124(A) (1961)\n6.126 S.H. Unger: Proc. IRE 46, 1744 (1958)\n6.127 J.H. Holland: Proc. WJCC, 259 (1960)\n6.128 W.T. Comfort: IBM Report No. 62-825-496 (1962)\n6.129 E.A. Feigenbaum, H.A. Simon: Proc. IFIP Congr. 1962, p. 177\n6.130 P.M. Davies: Proc. 1963 IEEE Pacific Compo Conf. (1963) p. 109\n6.131 P. Davies: \"Associative Processors\", IEEE Symp. on Search Memory, May 1964\n6.132 P.M. Davies: U.S. Patent No. 3,320,594, May 16, 1967\n6.133 E.V. Evreinov, V.G. Kosarev: Kibernetika 4, 3 (1963)\n6.134 R.G. Ewing, P.M. Davies: Proc. FJCC 25, 147 (1964)\n6.135 B. Hasbrouck, N.S. Prywes, D. Lefkovitz, N. Kornfield: Compo Command\nand Control Co., April 1965, AD-466 313\n6.136 R.G. Gall: \"Hybrid Associative Computer Study\", Vol. I, AD-489 929\n(Goodyear Aerospace Corp., 1966)\n6.137 R.G. Gall: \"Hybrid Associative Computer Study\", Vol. II, AD-489 930\n(Goodyear Aerospace Corp., 1966)\n6.138 R.G. Gall, D.E. Brotherton: \"Associative List Selector\", AD-802 993\n(Goodyear Aerospace Corp., 1966)\n6.139 D. L. Rohrbacher: \"Advanced Computer Organi zati on Study\", AD-631 870\nand AD-631 387 (April 1966)\n6.140 J.L. Cass: \"Organization and Applications of Associative File\nProcessors\", ONR/RADC Seminar on Associative Processing, ~1ay 1967\n6.141 T. Feng: \"An Associative Processor\"; Ph. D. Dissertation, Univ. of\nMichigan (1967)\n6.142 T. Feng: \"An Associative Processor\", Tech. Rept., Systems Engineering Lab., Univ. of Michigan, Dec. 1967\n6.143 T. Feng: \"An Associative Processor\", Michigan Univ. Rept.\nNo. 06920-17-T, AD-682 353 (Jan. 1969)\n6.144 T. Feng: Proc. Nat. Electron. Conf. XXIV, 257 (1968)\n6.145 Auerbach Publ. Inc.: TECH Note 1374-TR-500-1 (AD-679 227) (1968)\n6.146 W.A. Lea: NASA-TM-X1544, March 1968\n6.147 R.M. Lea: Radio and Electron. Eng. 46, 487 (1976)\n6.148 R.M. Lea: Comput. J. 21, 45 (1978)\n6.149 MIT Lincoln Lab.: Rept. No. ESD-TR-6890 (1968)\n6.150 H.H. Love: Hughes Aircraft Co., Rept. No. FR-69-11-487, AD-855 770\n(1969 )\n6.151 H.H. Love: Proc. Sagamore Comput. Conf. Parallel Process., Aug. 22-24,\n1973 (IEEE, New York 1973) p. 103\n6.152 P.M. Melliar-Smith: Proc. FJCC 1969, p. 201 350\n6.153\n6.154\n6.155\n6.156\n6.157\n6.158\n6.159\n6.160\n6.161\n6.162\n6.163\n6.164\n6.165\n6.166\n6.167\n6.168\n6.169\n6.170\n6.171\n6.172\n6.173\n6.174\n6.175\n6.176\n6.177\n6.178\n6.179\n6.180\n6.181\n6.182\n6.183\n6.184\n6.185\n6.186\n6.187\n6.188\nJ.E. Shore, F.A. Polkinghorn: NRL Rept. NRL-6961, Nov. 1969,\nAD-702 394\nW.S. Tuma: Goodyear Aerospace Corp., Rept. No. GER-14566, AD-862 134\n(1969)\nR.R. Kressler: Air Force Report No. AFAL-TR-70-142, Aug. 1970\nR.O. Berg, K.J. Thurber: NAECON '71 Record, p. 206 (1971)\nJ.E. Shore, T.L. Collins: Rept. of NRL Progress, p. 15, March 1972\nR.A. Urban: Nat. Electron. Conf. 1972, p. 318\nR.D. Arnold: Colorado Univ. Rept. CU CS 051 74, NSF GH 660, August 1974\nD.L. Baldauf: Mitre Corp., Bedford, Mass., MTR-2879, ESD-TR-74-199\n(AD-A003 414), Nov. 1974\nL.A. Gambino: Army Engineer. Topographic Labs., AD-A056 438, Jun. 1978\nG.J. Lipovski: Proc. 5th Annual Symp. Compo Archit. (IEEE, New York\n1978) p. 31\nS.Ya. Berkovich, Yu.Ya. Kochin, G.M. Lapir: Autom. Remote Control 35.\n1342 (1974)\nH.K. Resnick: California Univ., Livermore Lawrence Rad. Lab., Computer Inf. Center, Vol. 3, Publication No.6 (1975)\nL. Kerschberg, E.A. Ozkaharan, J.E.S. Pacheo: Proc. 2nd Int. Conf.\nSoftware Engineering, San Fransisco, Cal., 13-15 Oct., 1976\n(IEEE, New York 1976) p. 505\nC.Y. Hicks: ACM Compo Sci. Conf., 31 Jan.-2 Feb., 1977, Atlanta,\nGeorgia\nR.R. Seeber, A.B. Lindquist: Proc. AFIPS 1963 FJCC 24. 489 (1963)\nJ.S. Squire, S.M. Paleis: Proc. AFIPS 1963 SJCC, 395 (1963)\nR.S. Entner: \"The Advanced Avionic Digital Computer\", Symp. Parallel\nProcessor Systems, Tech. & Appl., Monterey, June 1969\nL.J. Koczela, G. Wang.: IEEE Electron. Comp., p. 520, June 1969\nG.J. Lipovski: Report R-424, Coordinated Sci. Lab., Univ. of Illinois,\nJuly 1969 (AD-692 195)\nC.C. Foster: Goodyear Aerospace Corp. Doc. GER-11772 (1964)\nM.J. Kroeger: Goodyear Aerospace Corp. Doc. GER-16378, RADC-TR-76-352\n(1976)\nZ.H. Glanz: Int. Electr. Electron. Conf. and Expos., 29 Sep.-1 Oct.,\n1975, Toronto, Canada\nB. Parhami, A. Avizienis: Symp. on Comput. Archit .. , Univ. of Florida,\nGainesville, p. 141 (1973)\nK.J. Thurber, P.C. Patton: COMPCON '72, p. 275 (1972)\nG.J. Nutt: Acta Infor. 6, 211 (1976)\nA.P. Kisylia: Illinois Univ. Rept. No. R-390, Aug. 1968 AD-675 310\nR.R. Linde, R. Gaten, T.F. Peng: Proc. AFIPS Nat. Compo Conf. 42,\n187 (1973)\nC.R. DeFiore: Datamation 16, 47 (1970)\nC.R. DeFiore, N.J. Stillman, P.B. Berra: Proc. ACM Nat. Conf.,\nAug. 3-5, 1971, p. 28\nV.L. Arlazarov, S.Ya. Berkovich, A.A. Leman, M.Z. Rosenfeld: Avtom.\nTelemekh. 12, 184 (1971)\nG. Salton: Commun. ACM 15, 658 (1972)\nC.R. DeFiore, P.B. Berra: Proc. AFIPS Conf. Nat. Compo Composition and Exposition 42, 181 (1973)\nC.R. DeFiore, P.B.Berra: IEEE Trans. C-23, 121 (1974)\nR. Moulder: Proc. Sagamore Comput. Conf. Parallel Process., Sagamore\nLake, N.Y. 1973 (IEEE, New York 1973) p. 161\nE.A. Ozkarahan, S.A. Schuster, K.C. Smith: Proc. AFIPS Nat. Comput. Conf. Expo. 44, 379 (1975)\nE.A. Ozkarahan, S.A. Schuster, K.C. Sevcik: ACM Trans. Database Syst. 2, 175 (1977) 6.189\n6.190\n6.191\n6.192\n6.193\n6.194\n6.195\n6.196\n6.197\n6.198\n6.199\n6.200\n6.201\n6.202\n6.203\n6.204\n6.205\n6.206\n6.207\n6.208\n6.209\n6.210\n6.211\n6.212\n6.213\n6.214\n6.215\n6.216\n6.217\n6.218\n6.219\n6.220\n6.221\n6.222\n6.223\n6.224\n6.225\n6.226\n6.227\n6.228\n6.229\nI.S. Chal~aya: Program. and Comput. Software 3, 61 (1977)\nR.E. Asratyan, V.T. Lysikov: Autom. Remote Control (USSR) 39, 755\n(1978)\n351\nR. Beaufils, J.P. Sansonnet: Euromicro J. (Netherlands) 4, 275 (1978)\nR.E. Birney, M.I. Davis, R.A. Hood: IBM Tech. Discl. Bull. 20,2972\n(1978)\nT. Ishikawa: AFIPS 1978 Nat. Compo Conf., 5-8 June, 1978, Anaheim, Cal.\nG.G. Langdon, Jr.: ACM Trans. Database Syst. 3, 148 (1978)\nS.A. Schuster, H.B. Nguyen, E.A. Ozkarahan, K.O. Smith: Proc. of the\n5th Ann. Symp. on Compo Archit., Palo Alto, 3-5 April, 1978\n(IEEE, New York 1978)\nJ.G. Dyke, R.M. Lea: Digital Process. 1, 89 (1975)\nV.A. Pronina, A.A. Chudin: Avtom. Telemekh. 8, 106 (1975)\nH.H. Love, J. Baer: Proc. 1977 Int. Conf. Parallel Processing\n(IEEE, New York 1977) p. 153\nR.M. Lea: Comput. J. 21, 45 (1978)\nR.H. Fuller, R.M. Bird: Proc. AFIPS 1965 FJCC 28, 105 (1965)\nC. Yang: Northwestern Univ. Tech. Rept. TR-66-103 (1966)\nC. Yang: \"Pattern Recognition by an Associative Memory\", Northwestern\nUniv. 1966 (unpublished paper) S.S. Yau, C.C. Yang: IEEE Trans. EC-15, 944 (1966)\nS.S. Yau, C.C. Yang: IEEE Trans. EC-15, 938 (1966)\nN.J. Stillman, C.R. DeFiore, P.R. Berra: Proc. AFIPS 1971 SJCC, 557\n(1971 )\nB. Kruse: IEEE Trans. C-22, 1075 (1973)\nE.C. Joseph, A. Kaplan: Proc. 6th Nat. MILECON, 1962, p. 255\nLibrascope Rept. Libi 6081, July 1966\nE.E. Eddey: Proc. NAECON '67, 39 (1967)\nE.E. Eddey: Proc. NAECON '70, 302 (1970)\nW.C. Meilander: Proc. NAECON '68, 57 (1968)\nL.E. Cannon: Ph. D. Thesis, Montana State Univ. (1969)\nA. Costanzo, J. Garrett: Proc. NAECON '69, 107 (1969)\nR.M. Bird: In Parallel Processor Systems, Technologies, and Applications,\ned. by L.C. Hobbs (Spartan Books, Washington, D.C. 1970) p. 107\nK.J. Thurber: Proc. AFIPS 1971 SJCC, 49 (1971)\nL.D. Wald: Proc. 1972 Sagamore Comput. Conf., Aug. 1972, (IEEE,\nNew York) p. 135\nL.D. Wald: Proc. AFIPS 1974 Nat. Comput. Conf., May 1974, p. 133\nE.E. Eddey, W.C. Meilander, T. Feng (ed.): Parallel Processing\n(Springer, Berlin, Heidelberg, New York 1975) p. 417\nC.L. Morefield: Proc. Annu. Allerton Conf. Circuit Syst. Theory, Monticello. Ill. Sep. 29- Oct. 1, 1976, p. 1074 (1976)\nH.G. Schmitz: Proc. ACM/ICST 15th Annu. Tech. Symp., June 17, 1976,\n4, 2338 (1976)\nA.K. Singhania: Proc. IMACS(AICA)-GI-Symp. Parallel Comput. March\n14-16, 1977, 5, 1071 (North Holland, Amsterdam 1977)\nS.M. Lamb, R. Vandersl ice: J. Acoust. Soc. 64, Sl, 573 (1978)\nG. Estrin, C.R. Viswanathan: J. ACM, Jan 1962, p. 41\nJ.H. Katz: In Parallel Processor Systems, Technologies, and Applications,\ned. by L.C. Hobbs (Spartan, Washington, D.C., 1970) p. 131\nP. Gilmore: Goodyear Aerospace Corp. Doc. GER-15260, June 1971\nP.B. Berra, E. Oliver: AD-A049 617/4SL, Syracuse Univ., Dept. of\nIndust. Eng. and Operations Res., Dec. 1977\nM.A. Wesley: IEEE Trans. AU-17, 162 (1969)\nYu.G. Naimark, G.M. Popova, I.V. Prangishvili: Avtom. Telemekh. No.4,\nApril 1972, p. 136\nA.J. Krygiel: Proc. 1976 Int. Conf. Parallel Process (IEEE, New York\n1976 ) 352\n6.230 P.A. Gilmore: Proc. AFIPS 1971 FJCC, 39, 411 (1971)\n6.231 W.F. Beausoleil, R.M. Chittenden, G.H. Ottaway: IBM Tech. Discl.\nBull. 20, 2770 (1977)\n6.232 W.C. Liles, J.C. Demmel, I.S. Reed, J.D. Mallett, L.E. Brennan:\nRept. No. TSC-PD-8525-1-Vol.-1, Apr. 1978, AD-A054 357\n6.233 W.C. Liles, J.C. Demmel, I.S. Reed, J.D. Mallett, L.E. Brennan:\nRept. No. TSC-PD-8525-1-Vol.-2, Apr. 1978, AD-A054 358\n6.234 M.E. Sherry: Amer. Document. Instit. 27th Ann. Meeting, 1964\n6.235 M.A. Wesley, S.K. Chang, J.H. Mommens: Proc. AFIPS 1972 FJCC, 461\n(1972)\n6.236 D.C. Gunderson: WESCON Tech. Papers (Session 9, 1966)\n6.237 J.P. Hayes: Univ. of Illinois, Comput. Lab. Rept. 227, June 1967\n6.238 J. Previte, E. Tippie: EMI-TM-67-1, Feb. 1967\n6.239 V.A. Orlando, P.B. Berra: Proc. AFIPS 1972 FJCC, 859 (1972)\n6.240 G.M. Popova, I.V. Prangishvili: Avtom. Telemekh. 1, 171 (1972)\n6.241 L.D. Wald, T.R. Armstrong, C.C. Huang, T.L. Saxton: RADC-TR-73-19\nFinal Tech. Rept., Feb. 1973\n6.242 W.T. Cheng, T.Y. Feng: Proc. 1974 Sagamore Comput. Conf. Parallel\nProcessing, Aug. 20-23, 1974 (IEEE, New York) p. 53\n6.243 W. Cheng, T. Feng: AD-A009 873, Syracuse Univ., Dept. of Electr.\nand Comput. Eng., March 1975 (RADC-TR-75-65)\n6.244 H.O. Welch: Proc. 1977 Int. Conf. Parallel Processing, ed. by J. Baer\np. 186 (IEEE, New York 1977)\n6.245 D.O. Marshall: Proc. 1977 Int. Conf. Parallel Processing, ed. by J.\nBaer p. 199 (IEEE, New York 1977)\n6.246 R. Napoli: Elettrotecnic. 65, 641 (1978)\n6.247 N.V. Findler: Cybernetica (Namur) 10, 229 (1967)\n6.248 N.V. Findler, W.R. McKinzie: Proc. Int. Joint Conf. Artificial\nIntelligence, May 1969, p. 259\n6.249 C.C. Foster: Univ. of Mass., Comput. Sci. Dept., TNCS-00023,\n(Dec. 1970)\n6.250 J.E. Shore: Rept. of NRL Prog., April 1972, p. 12\n6.251 B.F. Meyers: 8th Hawaii Int. Conf. Syst. Sci., 1975, p. 113\n6.252 W. Ash, E. Sibley: Univ. of Michigan, Tech. Rept. 5, June 1967\nAD-672 206\n6.253 W.L. Ash, E.H. Sibley: Proc. ACM 23rd Nat. Conf., 1968, p. 143\n6.254 W.L. Ash: Univ. of Michigan Rept. TR-17, May 1969 (AD-689 861)\n6.255 E.H. Sibley, R.W. Taylor, D.G. Gordon: Proc. AFIPS 1968 FJCC 33, 545\n(1968)\n6.256 P.O. Rovner, J.A. Feldman: MIT, Lincoln Lab. (AD-655 810), April 1967\n6.257 P.O. Rovner, J.A. Feldman: In Information Processing 68 (North-Holland,\nAmsterdam 1969) p. 579\n6.258 J.A. Feldman, P.O. Rovner: Stanford Univ. Rept. No. AI-Memo-66,\nAug. 1968 (AD-675 037)\n6.259 P.O. Rovner, D.A. Henderson, Jr.: Proc. Int. Joint Conf. Artificial\nIntelligence, May 1969, p. 9\n6.260 J.A. Feldman, J.R. Low, D.C. Swinehart, R.H. Taylor: Proc. AFIPS\n1972 FJCC 41, 1193 (1972)\n6.261 J.A. Feldman: Abst. of Tech. Repts. Comput. Sci. Dept. of Univ.\nRochester, TR9, Nov. 1976"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "All 22 reference pages were recovered from the publisher's first-edition back matter, including the full chapter-numbered sequence ending at 6.261. Every extracted source line through the final reference is present. The following subject index is excluded. Some reference numbers are detached by PDF layout extraction; no automatic citation links are inferred from that layout."
      ],
      "BibliographyEdition": "1980 first edition, DOI 10.1007/978-3-642-96552-4"
    },
    {
      "Slug": "hopfield",
      "Paper": "Neural networks and physical systems with emergent collective computational abilities",
      "AtlasYear": 1982,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://authors.library.caltech.edu/records/w41x7-8bn13/files/HOPpnas82.pdf?download=1",
      "PdfSha256": "1543E997FBBB49C3534DBFF1CEE399AB89C751BD1C98AC3CD3E7E5E22D5AF926",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 332,
          "EndLine": 360,
          "PdfPages": [
            5
          ],
          "Text": "1. Willows, A. 0. D., Dorsett, D. A. & Hoyle, G. (1973) J. Neurobiol 4, 207-237, 255-285.\n2. Kristan, W. B. (1980) in Information Processing in the Nervous System, eds. Pinsker, H. M. & Willis, W. D. (Raven, New York),\n241-261. 3. Knight, B. W. (1975) Lect. Math. Life Sci. 5, 111-144.\n4. Smith, D. R. & Davidson, C. H. (1962)J. Assoc. Comput. Mach.\n9, 268-279. 5. Harmon, L. D. (1964) in Neural Theory and Modeling, ed. Reiss,\nR. F. (Stanford Univ. Press, Stanford, CA), pp. 23-24. 6. Amari, S.-I. (1977) Bio. Cybern. 26, 175-185. 7. Amari, S.-I. & Akikazu, T. (1978) Biol Cybern. 29, 127-136. 8. Little, W. A. (1974) Math. Biosci. 19, 101-120.\n9. Marr, J. (1969) J. Physiol 202, 437-470. 10. Kohonen, T. (1980) Content Addressable Memories (Springer,\nNew York). 11. Palm, G. (1980) Biol Cybern. 36, 19-31. 12. McCulloch, W. S. & Pitts, W. (1943) BulL Math Biophys. 5,\n115-133.\n13. Minsky, M. & Papert, S. (1969) Perceptrons: An Introduction to Computational Geometry (MIT Press, Cambridge, MA).\n14. Rosenblatt, F. (1962) Principles of Perceptrons (Spartan, Washington, DC).\n15. Cooper, L. N. (1973) in Proceedings of the Nobel Symposium on Collective Properties of Physical Systems, eds. Lundqvist, B. & Lundqvist, S. (Academic, New York), 252-264.\n16. Cooper, L. N., Liberman, F. & Oja, E. (1979) Biol Cybern. 33,\n9-28.\n17. Longuet-Higgins, J. C. (1968) Proc. Roy. Soc. London Ser. B 171,\n327-334.\n18. Longuet-Higgins, J. C. (1968) Nature (London) 217, 104-105. 19. Kohonen, T. (1977) Associative Memory-A System-Theoretic\nApproach (Springer, New York). 20. Willwacher, G. (1976) Biol Cybern. 24, 181-198. 21. Anderson, J. A. (1977) Psych. Rev. 84, 413-451. 22. Perkel, D. H. & Bullock, T. H. (1969) Neurosci. Res. Symp.\nSumm. 3, 405-527. 23. John, E. R. (1972) Science 177, 850-864.\n24. Roney, K. J., Scheibel, A. B. & Shaw, G. L. (1979) Brain Res.\nRev. 1, 225-271.\n25. Little, W. A. & Shaw, G. L. (1978) Math. Biosci. 39, 281-289. 26. Shaw, G. L. & Roney, K. J. (1979) Phys. Rev. Lett. 74, 146-150. 27. Hebb, D. 0. (1949) The Organization of Behavior (Wiley, New\nYork).\n28. Eccles, J. G. (1953) The Neurophysiological Basis ofMind (Clar-\nendon, Oxford). 29. Kirkpatrick, S. & Sherrington, D. (1978) Phys. Rev. 17, 4384-4403. 30. Mountcastle, V. B. (1978) in The Mindful Brain, eds. Edelman,\nG. M. & Mountcastle, V. B. (MIT Press, Cambridge, MA), pp.\n36-41.\n31. Goldman, P. S. & Nauta, W. J. H. (1977) Brain Res. 122,\n393-413. 32. Kandel, E. R. (1979) Sci. Am. 241, 61-70."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "backpropagation",
      "Paper": "Learning representations by back-propagating errors",
      "AtlasYear": 1986,
      "Status": "indexed",
      "Method": "publisher-reference-section",
      "SourceUrl": "https://www.nature.com/articles/323533a0",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "Publisher references",
          "StartLine": 1,
          "EndLine": 4,
          "PdfPages": [],
          "Text": "1. Rosenblatt, F. Principles of Neurodynamics (Spartan, Washington, DC, 1961).\n2. Minsky, M. L. & Papert, S. Perceptrons (MIT, Cambridge, 1969).\n3. Le Cun, Y. Proc. Cognitiva 85, 599–604 (1985).\n4. Rumelhart, D. E., Hinton, G. E. & Williams, R. J. in Parallel Distributed Processing: Explorations in the Microstructure of Cognition. Vol. 1: Foundations (eds Rumelhart, D. E. & McClelland, J. L.) 318–362 (MIT, Cambridge, 1986)."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference section extracted from the linked source. Only reviewed matches to existing Atlas works become citation links; extraction may retain typographic or column-order artifacts."
      ]
    },
    {
      "Slug": "dynamic-error-propagation",
      "Paper": "The Utility Driven Dynamic Error Propagation Network",
      "AtlasYear": 1987,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://gwern.net/doc/ai/nn/rnn/1987-robinson.pdf",
      "PdfSha256": "D20459C799BAC3D9A406CC07F8D351F4EB1C903D1B84123EBEF81BE54E240C30",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 735,
          "EndLine": 788,
          "PdfPages": [
            23
          ],
          "Text": "\nAckley, D. H., Hinton, G. E., and Sejnowski, T. J. (1985). A learning algorithm for\n\nBoltzmann machines. Journal of Cognitive Science, 9, 147-169.\n\nCottrell, G. W., Munro, P., and Zipser, D. (Febuary 1986). Image Compression by Back\n\nPropagation: An Example of Existential Programming. ICS Report 8702, Institute for\n\nCognitive Science, University of California, San Diego.\n\n.\n\nDennett, D. C. (1984). Elbow Room: The varieties of free will worth wanting. The Clarendon\n\nPress, Oxford.\n\nElman, J. L. and Zipser, D. (1987). Learning the Hidden Structure of Speech. ICS Re-\nport 8701, University of California, San Diego. Hofstader, D. R. (1979). Godel, Escher, Bach: An eternal golden braid. The Harvester Press,\nHassocks, Sussex.\n\nHopfield, J. J. (1982). Neural networks and physical systems with emergent collective\ncomputationalabilities. Proceedings of the National Academy of Science U.S.A., 79, 2554—\n2558.\n\nJacobs, O. L. R. (1974). Introduction to Control Theory. Clarendon Press, Oxford.\nJordan, M. I. (May 1986). Serial Order: A Parallel Distributed Processing Approach. ICS\nReport 8604, Institute for Cognitive Science, University of California, San Diego.\nKuffler, S. W., Nicholls, J. G., and Martin, A. R. (1984). From Neuron to Brain: A Cellular Approach to the Function of the Nervous System. Sinauer Associates Inc., Sunderland,\nMA, second edition.\n\nLindsay, P. H. and Norman, D. A. (1977). Human Information Processing: An Introduction\nto Psychology. Academic Press, Inc., Orlando, Florida, second edition. Minsky, M. and Papert, S. (1969). Perceptrons: An Introduction to Computational Geometry.\nMIT Press, Cambridge, MA.\n\nPearlmutter, B. A. and Hinton, G. E. (1986). G-maximization: An unsupervised learning procedure for discovering regularities. In Proceedings of the Conference on ‘Neural Networks for Computing’, American Institute of Physics.\nPoggio, T. and Koch, C. (1987). Synapses that compute motion. Scientific American, May,\n42-48.\n\nPrager, R. W., Harrison, T. D., and Fallside, F. (1986). Boltzmann machines for speech\nrecognition. Computer Speech and Language, 1, 3-27.\nRobinson, A. J. (1986). Speech Recognition with Associative Networks. M.Phil Computer Speech and Language Processing thesis , Cambridge University Engineering Depart-\nment.\n\nRosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organisation in the brain. Psychological Review, 65, 386-408.\n\nRosenblatt, F. (1962). Principles of Neurodynamics. Spartan, New York.\nRumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986). Learning internal representations by error propagation. In Parallel Distributed Processing: Explorations in the\nMicrostructure of Cognition. Vol. 1: Foundations. (eds. D. E. Rumelhart and J. L. McClel-\nland), Bradford Books/MIT Press, Cambridge, MA. Rumelhart, D. E. and McClelland, J. L. (1986). Parallel Distributed Processing: Explorations\nim the Microstructure of Cognition. Vol. 1: Foundations. MIT Press, Cambridge, MA. Tank, D. W. and Hopfield, J. J. (1987). Neural computation by concentrating information\nin time. Proceedings of the National Academy of Science U.S.A.. 84, 1896-1900. Watrous, R. L., Shastri, L., and Waibel. A. H. (1987). Learned phonetic discrimina-\ntion using connectionist networks. In Proceedings of the European Conference on Speech\nTechnology (eds. J. Laver and M. A. Jack), CEP Consultants Ltd, Edinburgh."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "finding-structure-in-time",
      "Paper": "Finding Structure in Time",
      "AtlasYear": 1988,
      "Status": "indexed",
      "Method": "windows-ocr-author-designated-journal-version",
      "SourceUrl": "https://cogsci.ucsd.edu/~rik/courses/readings/elman90-fsit.pdf",
      "PdfSha256": "B019195E44723933645DBF2DF267C1E4E39FD5DDEE341FB705FF127F610B609B",
      "Sections": [
        {
          "Section": "References, printed pages 209-211",
          "PdfPages": [
            31,
            32,
            33
          ],
          "Text": "REFERENCES\nBerwick, R.C.. & Weinberg, A.S. (1984). The grammatical of linguistic performance.\nCambridge, MA: MIT Pres.\nChomsky, N. (1957). Syntactic structures. The Hague: Moutin.\nChomsky, N. (1965). Aspects of the theory of syntax. Cambridge. MA: MIT Pres.\nCottrell, G.W., Munro. P.W., & Zipser. D. (1987). Image comprasion by back propagation:\nA demonstration of extensional programming. In N.E. Sharkey (Ed.), Adpanæs in\ncognitive science (Vol. 2). Chichester. England: Ellis Horwood.\nElman. J.L. (1989). Structurd representations and connectionist models. (CRL Tech. Rep.\nNo. 8901). San Diego: University of California, Center for Research in Language.\nElman, J.L., & Zipser, D. (1988). Discovering the hidden structure of speech. Journal of the\nAcoustical Society of America, U, 1615-1626.\nFodor. J & Pylyshyn. Z. (1988). Connectionism and cognitive architecture: A critical analysis.\nIn S. Pinker & J. Mehler (Eds.), Connections and symbols (pp. 3-71). Cambridge, MA:\nMIT Press.\nFowler. C. (1977). Timing control in speeh production. Bloomington. IN: Indiana University\nLinguistics Club.\nFowler , C. (1980). Coarticulation and theories of extrinsic timing control. Journal of Phonetics.\n8. 113-133.\nFrazier, L. , & Fodor, J.D. (1978). The sausage machine: A new two-stage parsing model.\nCognition. 6. 291-325.\nGreenberg, J.H. (1963). Universals of language. Cambridge, MA: MIT Press.\nGrosjean, F. (1980). Spokel word recognition process— and the gating paradigm. Perception\n& Psychophysics. 28, 267-283.\nHanson, S.J., & Keg), J. (1987). Parsnip: A connectionist network that learns natural language\ngrammar from exposure to natural language sentence. Ninth A naual Confeænce of the\nCognitive Science Society, Seattle, Washington. Hillsdale, NJ: Erlbaum.\nHinton, G.E., McClelland, J.L., & Rumelhart, D.E. (1986). Distributed repreentations. In\nD.E. Rumelhart & J.L. McClelland (Eds.), hm'lel distributed procsing:\n210\nELMAN\n(ions in the microstructure of cognition (Vol. l, pp. 77-109). Cambridge, MA: MIT\nPress.\nJordan, M.I. (1986). Serial order: A parallel distributed processing approach (Tech. Rep. No.\n804). San Diego: University of California, Institute for Cognitive Science.\nJordan, M.I., & Rosenbaum, D.A. (1988). Action (Tech. Rep. No. 88-26). Amherst: Univer-\nsity of Massachusetts, Department of Computer Science.\nKelso, J.A.S., Saltzman, E., & Tuner, B. (1986). The dynamical theory of speech production:\nData and theory. Journal of Phonetics, 14, 29-60.\nLashley, KS. (1951). The problem of serial order in behavior. In L.A. Jeffress (Ed.), Cerebral\nmechanisms in behavior. New York: Wiley.\nLehman, W.P. (1962). Historical linguistics: An introduction. New York: Holt, Rinehart, and\nWinston.\nMacNeilage, P.F. (1970). Motor control of serial ordering of speech. Psychological Review.\n77, 182-196.\nMacWhinney, B. (1978). The acquisition of morphophonology. Monographs of the Society for\nResearch in Child Development,' 43, (Serial No. l).\nMarcus, M. (1980). A theory of syntactic recognition for natural language. Cambridge, MA:\nMIT Press.\nMarslen-Wilson, W., & Tyler, L.K. (1980). The temporal structure of spoken language under-\nstanding. Cognition, 8, I-II.\nPineda, F.J. (1988). Generalization of back propagation to recurrent and higher order neural\nnetworks. In D.Z. Anderson (Ed.), Neural information processing systems. New York:\nAmerican Institute of Physics.\nPinker, S. (1984). Language learnability and language development. Cambridge, MA: Harvard\nUniversity Press.\nRumelhart, D.E., Hinton, G.E., & Williams, R.J. (1986). Learning internal representations by\nerror propagation. In D.E. Rumelhart & J.L. McClelland (Eds.), Parallel distributed\nprocessing: Explorations in the microstructure ofcognition (Vol. l, pp. 318-362). Cam-\nbridge, MA: MIT Press.\nSalasoo, A., & Pisoni, D.B. (1985). Interaction of knowledge sources in spoken word identifi-\ncation. Journal of Memory and Language, 24, 210-231.\nSaltzman, E. , & Kelso, J. A.S. (1987). Skilled actions: A task dynamic approach. Psychological\nReview, 94, 84-106.\nSchwanennugel, P.J., & Shoben, E.J. (1985). The influence of sentence constraint on the\nscope of facilitation for upcoming words. Journal of Memory and Language, 24,\n232-252.\nSejnowski, T.J ., & Rosenberg, C.R. (1987). Parallel networks that learn to pronounce English\ntext. Complex Systems, l, 145-168.\nServan-Schreiber, D., Cleeremans, A. , & McClelland, J. L. (1988). Encoding sequential struc-\nture in simple recurrent networks (CMU Tech. Rep. No. CMU-CS-88-183). Pittsburgh,\nPA: Carnegie-Mellon University, Computer Science Department.\nSmolensky, P. (1987). On variable binding and the representation of symbolic structures in\nconnectionist systems (Tech. Rep. No. CU-CS-355-87). Boulder, CO: University of\nColorado, Department of Computer Science.\nSmolensky, P. (1988). On the proper treatment of connectionism. The Behavioral and Brain\nSciences, 11.\nStornetta, W.S., Hogg, T. , & Huberman, B.A. (1987). A dynamical approach to temporal\npattern processing. Proceedings of the IEEE Conference on Neural Idormation Pro-\ncessing Systems. Denver, CO.\nSwinney, D. (1979). Lexical access during sentence comprehension: (Re)consideration Of con-\ntext effects. Journal of Verbal Learning and Verbal Behavior, 6, 645-659.\nFINDING STRUCTURE IN TIME\n211\nTabossi, P. (1988). Effects of context on the immediate interpretation of unambiguous nouns.\nJournal of Experimental Psychology: Learning, Memory, and Cognition, 14, 153-162.\nTabossi, P. , Colombo, L. , & Job, R. (1987). Accessing lexical ambiguity: Effects of context\nand dominance. Psychological Research, 49, 161-167.\nTank, D.W., & Hopfield, J. J. (1987, June). Neural computation by concentrating information\nin time. Proceedings of the IEEE International Conference on Neural Networks. San\nDiego, CA.\nVan Gelder, T.J. (in press). Compositionality: Variations on a classical theme. Cognitive\nScience.\nWaibel, A. , Hanazawa, T., Hinton, G. , Shikano, K. , & Lang, K. (1987). Phoneme recognition\nusing time-delay neural networks (ATR Tech. Rep. TR-I-0006). Japan: ATR Interpret-\ning Telephony Research Laboratories.\nWatrous, R.L., & Shastri, L. (1987). Learning phonetic features using connectionist networks:\nAn experiment in speech recognition. Proceedings of the IEEE International Confer-\nence on Neural Networks. San Diego, CA.\nWilliams, R.J., & Zipser, D. (1988). A learning algorithm for continually runningfully recur-\nrent neural networks (Tech. Rep. No. 8805). San Diego: University of California, Insti-\ntute for Cognitive Science."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "The author's publication archive states that the 1988 technical report is no longer available and was superseded by the 1990 journal paper: https://jeffelman.ucsd.edu/research/publications/. The complete bibliography of that author-designated replacement is indexed here. All three reference pages were processed with Windows OCR; image and text hashes are retained. This does not establish the bibliography of the unavailable 1988 report. The Atlas chronology continues to date the first report to 1988."
      ],
      "BibliographyEdition": "1990 Cognitive Science journal version, DOI 10.1207/s15516709cog1402_1; replacement for the 1988 report",
      "OcrPages": [
        {
          "PdfPage": 31,
          "Text": "FINDING STRUCTURE IN TIME\nactivation pattern time series, and then to construct phase state portraits of\nthe most significant principal components (Elman, 1989).\nAnother question of interest is what is the memory capacity of such net-\nworks. The results reported here suggest that these networks have consider-\nable representational power; but more systematic analysis using better defined\ntasks is clearly desirable. Experiments are currently underway using sequences\ngenerated by finite state automata of various types; these devices are rela-\ntively well understood, and their memory requirements may be precisely\ncontrolled (Servan-Schreiber et al., 1988).\nOne of the things which feedforward PDP models have shown is that\nsimple networks are capable of discovering useful and interesting internal\nrepresentations of many static tasks. Or put the other way around: Rich\nrepresentations are implicit in many tasks. However, many of the most in-\nteresting human behaviors have a serial component. What is exciting about\nthe present results is that they suggest that the inductive power of the PDP\napproach can be used to discover structure and representations in tasks\nwhich unfold over time.\nREFERENCES\nBerwick, R.C.. & Weinberg, A.S. (1984). The grammatical of linguistic performance.\nCambridge, MA: MIT Pres.\nChomsky, N. (1957). Syntactic structures. The Hague: Moutin.\nChomsky, N. (1965). Aspects of the theory of syntax. Cambridge. MA: MIT Pres.\nCottrell, G.W., Munro. P.W., & Zipser. D. (1987). Image comprasion by back propagation:\nA demonstration of extensional programming. In N.E. Sharkey (Ed.), Adpanæs in\ncognitive science (Vol. 2). Chichester. England: Ellis Horwood.\nElman. J.L. (1989). Structurd representations and connectionist models. (CRL Tech. Rep.\nNo. 8901). San Diego: University of California, Center for Research in Language.\nElman, J.L., & Zipser, D. (1988). Discovering the hidden structure of speech. Journal of the\nAcoustical Society of America, U, 1615-1626.\nFodor. J & Pylyshyn. Z. (1988). Connectionism and cognitive architecture: A critical analysis.\nIn S. Pinker & J. Mehler (Eds.), Connections and symbols (pp. 3-71). Cambridge, MA:\nMIT Press.\nFowler. C. (1977). Timing control in speeh production. Bloomington. IN: Indiana University\nLinguistics Club.\nFowler , C. (1980). Coarticulation and theories of extrinsic timing control. Journal of Phonetics.\n8. 113-133.\nFrazier, L. , & Fodor, J.D. (1978). The sausage machine: A new two-stage parsing model.\nCognition. 6. 291-325.\nGreenberg, J.H. (1963). Universals of language. Cambridge, MA: MIT Press.\nGrosjean, F. (1980). Spokel word recognition process— and the gating paradigm. Perception\n& Psychophysics. 28, 267-283.\nHanson, S.J., & Keg), J. (1987). Parsnip: A connectionist network that learns natural language\ngrammar from exposure to natural language sentence. Ninth A naual Confeænce of the\nCognitive Science Society, Seattle, Washington. Hillsdale, NJ: Erlbaum.\nHinton, G.E., McClelland, J.L., & Rumelhart, D.E. (1986). Distributed repreentations. In\nD.E. Rumelhart & J.L. McClelland (Eds.), hm'lel distributed procsing:",
          "Receipt": {
            "TextSha256": "BB68EB26C844328584C8E65070AEB87A930F1985DD8EB5BB8EA71AA6C7BA8F2E",
            "ProcessedAtUtc": "2026-09-17T00:04:30.1934283Z",
            "ImageSha256": "1B8B5F77B400BBED254894904FD4432C0DFCD9FCA63C01E93A2E80AD41C224C3",
            "Language": "en-US",
            "Lines": 47,
            "Engine": "Windows.Media.Ocr",
            "Output": "elman-refs-000031.ocr.txt",
            "Image": "elman-refs-000031.png"
          }
        },
        {
          "PdfPage": 32,
          "Text": "210\nELMAN\n(ions in the microstructure of cognition (Vol. l, pp. 77-109). Cambridge, MA: MIT\nPress.\nJordan, M.I. (1986). Serial order: A parallel distributed processing approach (Tech. Rep. No.\n804). San Diego: University of California, Institute for Cognitive Science.\nJordan, M.I., & Rosenbaum, D.A. (1988). Action (Tech. Rep. No. 88-26). Amherst: Univer-\nsity of Massachusetts, Department of Computer Science.\nKelso, J.A.S., Saltzman, E., & Tuner, B. (1986). The dynamical theory of speech production:\nData and theory. Journal of Phonetics, 14, 29-60.\nLashley, KS. (1951). The problem of serial order in behavior. In L.A. Jeffress (Ed.), Cerebral\nmechanisms in behavior. New York: Wiley.\nLehman, W.P. (1962). Historical linguistics: An introduction. New York: Holt, Rinehart, and\nWinston.\nMacNeilage, P.F. (1970). Motor control of serial ordering of speech. Psychological Review.\n77, 182-196.\nMacWhinney, B. (1978). The acquisition of morphophonology. Monographs of the Society for\nResearch in Child Development,' 43, (Serial No. l).\nMarcus, M. (1980). A theory of syntactic recognition for natural language. Cambridge, MA:\nMIT Press.\nMarslen-Wilson, W., & Tyler, L.K. (1980). The temporal structure of spoken language under-\nstanding. Cognition, 8, I-II.\nPineda, F.J. (1988). Generalization of back propagation to recurrent and higher order neural\nnetworks. In D.Z. Anderson (Ed.), Neural information processing systems. New York:\nAmerican Institute of Physics.\nPinker, S. (1984). Language learnability and language development. Cambridge, MA: Harvard\nUniversity Press.\nRumelhart, D.E., Hinton, G.E., & Williams, R.J. (1986). Learning internal representations by\nerror propagation. In D.E. Rumelhart & J.L. McClelland (Eds.), Parallel distributed\nprocessing: Explorations in the microstructure ofcognition (Vol. l, pp. 318-362). Cam-\nbridge, MA: MIT Press.\nSalasoo, A., & Pisoni, D.B. (1985). Interaction of knowledge sources in spoken word identifi-\ncation. Journal of Memory and Language, 24, 210-231.\nSaltzman, E. , & Kelso, J. A.S. (1987). Skilled actions: A task dynamic approach. Psychological\nReview, 94, 84-106.\nSchwanennugel, P.J., & Shoben, E.J. (1985). The influence of sentence constraint on the\nscope of facilitation for upcoming words. Journal of Memory and Language, 24,\n232-252.\nSejnowski, T.J ., & Rosenberg, C.R. (1987). Parallel networks that learn to pronounce English\ntext. Complex Systems, l, 145-168.\nServan-Schreiber, D., Cleeremans, A. , & McClelland, J. L. (1988). Encoding sequential struc-\nture in simple recurrent networks (CMU Tech. Rep. No. CMU-CS-88-183). Pittsburgh,\nPA: Carnegie-Mellon University, Computer Science Department.\nSmolensky, P. (1987). On variable binding and the representation of symbolic structures in\nconnectionist systems (Tech. Rep. No. CU-CS-355-87). Boulder, CO: University of\nColorado, Department of Computer Science.\nSmolensky, P. (1988). On the proper treatment of connectionism. The Behavioral and Brain\nSciences, 11.\nStornetta, W.S., Hogg, T. , & Huberman, B.A. (1987). A dynamical approach to temporal\npattern processing. Proceedings of the IEEE Conference on Neural Idormation Pro-\ncessing Systems. Denver, CO.\nSwinney, D. (1979). Lexical access during sentence comprehension: (Re)consideration Of con-\ntext effects. Journal of Verbal Learning and Verbal Behavior, 6, 645-659.",
          "Receipt": {
            "TextSha256": "179A3ED7F608BACF847F95F82DC44F35C4225CA8E789FFAB91737C3345E69213",
            "ProcessedAtUtc": "2026-09-17T00:04:30.385455Z",
            "ImageSha256": "266F4D8B517DBD48D48344DD4CC2112EAD6E2526582443D19E9CF8A3281772B1",
            "Language": "en-US",
            "Lines": 53,
            "Engine": "Windows.Media.Ocr",
            "Output": "elman-refs-000032.ocr.txt",
            "Image": "elman-refs-000032.png"
          }
        },
        {
          "PdfPage": 33,
          "Text": "FINDING STRUCTURE IN TIME\n211\nTabossi, P. (1988). Effects of context on the immediate interpretation of unambiguous nouns.\nJournal of Experimental Psychology: Learning, Memory, and Cognition, 14, 153-162.\nTabossi, P. , Colombo, L. , & Job, R. (1987). Accessing lexical ambiguity: Effects of context\nand dominance. Psychological Research, 49, 161-167.\nTank, D.W., & Hopfield, J. J. (1987, June). Neural computation by concentrating information\nin time. Proceedings of the IEEE International Conference on Neural Networks. San\nDiego, CA.\nVan Gelder, T.J. (in press). Compositionality: Variations on a classical theme. Cognitive\nScience.\nWaibel, A. , Hanazawa, T., Hinton, G. , Shikano, K. , & Lang, K. (1987). Phoneme recognition\nusing time-delay neural networks (ATR Tech. Rep. TR-I-0006). Japan: ATR Interpret-\ning Telephony Research Laboratories.\nWatrous, R.L., & Shastri, L. (1987). Learning phonetic features using connectionist networks:\nAn experiment in speech recognition. Proceedings of the IEEE International Confer-\nence on Neural Networks. San Diego, CA.\nWilliams, R.J., & Zipser, D. (1988). A learning algorithm for continually runningfully recur-\nrent neural networks (Tech. Rep. No. 8805). San Diego: University of California, Insti-\ntute for Cognitive Science.",
          "Receipt": {
            "TextSha256": "E5CAE9503FFE2F02EE0B6CA406CA7BBEFC582E34733A3F73616182C08CF9345E",
            "ProcessedAtUtc": "2026-09-17T00:04:30.4812752Z",
            "ImageSha256": "912D28390E5EBCEA80DEAEE1CF1571C52C9432900B4066CCB44167AD99B4BC93",
            "Language": "en-US",
            "Lines": 20,
            "Engine": "Windows.Media.Ocr",
            "Output": "elman-refs-000033.ocr.txt",
            "Image": "elman-refs-000033.png"
          }
        }
      ]
    },
    {
      "Slug": "universal-approximation",
      "Paper": "Multilayer feedforward networks are universal approximators",
      "AtlasYear": 1989,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://cognitivemedium.com/magic_paper/assets/Hornik.pdf",
      "PdfSha256": "C4DC3A119F37E717C3E2628F57F6FC7E798B514477520664AE68E7DAB25748C9",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 808,
          "EndLine": 812,
          "PdfPages": [
            8
          ],
          "Text": "Billingsley, P. (1979). Probability and measure. New York: Wiley. Cybenko, G. (1988). Approximation by superpositions of a sig-\nmoidaifunction (Tech. Rep. No. 856). Urbana, IL: University of Illinois Urbana-Champaign Department of Electrical and Computer Engineering. Dugundji, J. (1966). Topology. Boston: Allyn and Bacon, Inc. Gallant, A. R., &White, H. (1988). There exists a neural network that does not make avoidable mistables. In IEEE Second International Conference on Neural Networks (pp. 1:657-X164). San Diego: SOS Printing. Grenander, U. (1981). Abstract inference. New York: Wiley. Halmos, P. R. (1974). Measure theory. New York: Springer-Verlag. Hecht-Nielsen, R. (1987). Kolmogorov’s mapping neural network existence theorem. In IEEE First International Conference on Neural Networks (pp. III:ll-14). San Diego: SOS Printing. Hecht-Nielsen, R. (1989). Theory of the back propagation neural network. In Proceedings of the International Joint Conference on Neural Networks (pp. 1593408). San Diego: SOS Printing. Hornik, K., Stinchcombe, M., & White, H. (1988). Multilayer feedforward networks are universal approximators (Discussion Paper 88-45). San Diego, CA: Department of Economics, University of California, San Diego. IEEE First International Conference on Neural Networks (1987). M. Caudill and C. Butler (Eds.). San Diego: SOS Printing. IEEE Second International Conference on Neural Networks (1988). San Diego: SOS Printing. Irie, B., & Miyake, S. (1988). Capabilities of three layer perceptrans. In IEEE Second International Conference on Neural Networks (pp. 1641-648). San Diego: SOS Printing Kolmogorov, A. N. (1957). On the representation of continuous\n\nK. Hornik, M. Stinchcornhe, und H. White\nfunctions of many variables by superposition of continuous functions of one variable and addition. Doklady Akademii Nauk SSR, 114, 953-956. Kolmogorov, A. N.. & Tihomirov. V. M. (1961). E-entropy and e-capacity of sets in functional spaces. American Mathematical Society Translations, 2(17), 277-364. Lapedes. A., & Farber. R. (1988). How neural networks work (Tech. Rep. LA-UR-88-418). Los Alamos. NM: Los Alamos National Laboratory. le Cun, Y. (1987). Modeles connexiontstes de l’upprentissage. IIese de Doctorat, Universite Pierre et Marie Curie. Lorentz, G. G. (1976). The thirteenth problem of Hilbert. In F E. Browder (Ed.), Proceedings of Symposia in Pure Mathematics (Vol. 28, pp. 419-430). Providence. RI: American Mathematical Society. Maxwell. T., Giles, G. L.. Lee, Y. C., & Chen, H. H. (1986). Nonlinear dynamics of artificial neural systems. In J. Denker (Ed.), Neural networks for computing. New York: American Institute of Physics. Minsky, M., & Papert, S. (1969). Perceprrons. Cambridge: MIT Press. Rudin. W. (1964). Principles of mathematrcal analysts. New York: McGraw-Hill. Severini, J. A., & Wang, W. H. (1987). C.‘onvergence rates of maximum likelihood and related estimates in general parameter vpaces (Working Paper). Chicago. IL: l!niversity of Chicago Department of Statistics. Stinchcombe, M., & White, H. (1989). Universal approximation using feedforward networks with non-sigmoid hidden layer activation functions. In Proceedings off the international Joint Conference on Neural Networks (pp. 1:613-618). San Diego: SOS Printing. White, H. (1988a). The case for conceptual and operational separation of network architectures and learning mechanisms (Discussion Paper 88-21). San Diego, CA: Department of Economics, University of California, San Diego. White, H. (1988b). Multilayer feedforward networks can learn arbitrary mappings: Connectionist nonparametric regression with automatic and semi-automatic determination of network complexity (Discussion Paper). San Diego, CA: Department of Economics. University of California, San Diego. White, H., & Wooldridge, J. M. (in press). Some results for sieve estimation with dependent observations. In W. Barnett. J. Powell. & G. Tauchen (Eds.), Nonparametric and semi-parametric methods in econometrics and statistus. New York: Cambridge University Press. Williams, R. J. (1986). The logic of activation functions. In D. E. Rumelhart & J. L. McClelland (Eds.), Parallel distributed processing: Explorations in the microstructures of cognition (Vol. 1, pp. 423-443). Cambridge: MIT Press."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "finite-state-recurrent-networks",
      "Paper": "Finite State Automata and Simple Recurrent Networks",
      "AtlasYear": 1989,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://consciousbrain.ulb.ac.be/uploads/2015/11/89-nc.pdf",
      "PdfSha256": "D1B976341D4C74DA85A23998BF54099B7E74B7E3C8F74F38961823E785EE610A",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 139,
          "EndLine": 147,
          "PdfPages": [
            10
          ],
          "Text": "Elman, J.L. 1988. Finding Structure in Time. CRL Tech. Rep. 9901. Center for Research in Language, University of California, San Diego, CA.\nMcClelland, J.L. 1988. The case for interactionism in language processing. In Attention and Performaiice XII, M. Coltheart, ed. Erlbaum, London.\nReber, A.S. 1967. Implicit learning of artificial grammars. I. Verbal Leartiing\nVerbal Behau. 5, 855-863. Rumelhart, D.E., Hinton, G.E., and Williams, R.J. 1986. Learning internal rep-\nresentations by backpropagating errors. Nature (London) 323,533-536. Sejnowski,T.J.,and Rosenberg, C. 1986. NETtalk: A Parallel Network That Learns\nto Read Aloud. Tech. Rep. JHU-EECS-86-01,Johns Hopkins University. Servan-Schreiber, D., Cleeremans, A,, and McClelland, J.L. 1988. Learnitig Se-\nquential Structure in Sirnple Recirrretit Netzuorks. Tech. Rep. CMU-CS-183, Carnegie-Mellon University. Williams, R.J., and Zipser, D. 1988. A Learning Algorithm for Cotititiually Running Fully Recurrent Neural Netzuorks. ICS Tech. Rep. 8805. Institute for Cognitive Science, University of California, San Diego, CA."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference text retains scan, spelling, line-break and column-order artifacts. Matching does not rely on publication year alone."
      ]
    },
    {
      "Slug": "focused-backpropagation",
      "Paper": "A Focused Backpropagation Algorithm for Temporal Pattern Recognition",
      "AtlasYear": 1989,
      "Status": "indexed",
      "Method": "windows-ocr-with-reviewed-citation-metadata",
      "SourceUrl": "https://home.cs.colorado.edu/~mozer/Research/Selected%20Publications/reprints/Mozer1989.pdf",
      "PdfSha256": "5F865546F4FF9C13674B9565970C4D9D167E37F56B0727FDC256BA5E8F7087E3",
      "Sections": [
        {
          "Section": "Bibliography",
          "StartLine": 1,
          "EndLine": 33,
          "PdfPages": [
            16,
            17
          ],
          "Text": "1. D.E. Rumelhart and J.L. McClelland. On learning the past tenses of English verbs. Parallel Distributed Processing, volume II, 1986, 216-271.\n2. T.J. Sejnowski and C.R. Rosenberg. Parallel networks that learn to pronounce English text. Complex Systems 1 (1987), 145-168.\n3. G. Hinton. Learning distributed representations of concepts. Proceedings of the Eighth Annual Conference of the Cognitive Science Society (1987), 1-12.\n4. P. Smolensky. Schema selection and stochastic inference in modular environments. Proceedings AAAI-83 (1983), 109-113.\n5. J. Freyd. Dynamic mental representations. Psychological Review 94 (1987), 427-438.\n6. J.L. Elman and J.L. McClelland. Exploiting lawful variability in the speech wave. In Invariance and variability in speech processes (1986), 360-380.\n7. J.L. Elman and D. Zipser. Learning the hidden structure of speech. Journal of the Acoustical Society of America, in press.\n8. T.K. Landauer, C.A. Kamm and S. Singhal. Teaching a minimally structured back propagation network to recognize speech. Proceedings of the Ninth Annual Conference of the Cognitive Science Society (1987), 531-536.\n9. A. Lapedes and R. Farber. Nonlinear signal processing using neural networks. Report LA-UR-87-2662, Los Alamos, 1987.\n10. J.L. McClelland and J.L. Elman. Interactive processes in speech perception: The TRACE model. Parallel Distributed Processing, volume II (1986), 58-121.\n11. D.C. Plaut, S. Nowlan and G.E. Hinton. Experiments on learning by back propagation. Technical Report CMU, Carnegie-Mellon University, 1986.\n12. D. Tank and J. Hopfield. Proceedings of the National Academy of Sciences 84 (1987), 1896. No article title supplied in the reference.\n13. A. Waibel, T. Hanazawa, G. Hinton, K. Shikano and K. Lang. Phoneme recognition using time-delay neural networks. Technical Report 1-0006, ATR Interpreting Telephony Research Laboratories, 1987.\n14. G. Hinton. Connectionist learning procedures. Artificial Intelligence (1988), in press.\n15. K. Lang. Connectionist speech recognition. Unpublished Ph.D. thesis proposal, Carnegie-Mellon University, 1987.\n16. J.L. Elman. Finding structure in time. CRL Technical Report 8801, University of California, San Diego, Center for Research in Language, 1988.\n17. W.S. Stornetta, T. Hogg and B.A. Huberman. A dynamical approach to temporal pattern processing. Proceedings of the IEEE Conference on Neural Information Processing Systems, 1987.\n18. R.L. Watrous and L. Shastri. Learning acoustic features from speech data using connectionist networks. Proceedings of the Ninth Annual Conference of the Cognitive Science Society (1987), 518-530.\n19. M.I. Jordan. Attractor dynamics and parallelism in a connectionist sequential machine. Proceedings of the Eighth Annual Conference of the Cognitive Science Society (1987), 531-546.\n20. D.E. Rumelhart, G.E. Hinton and R.J. Williams. Learning internal representations by error propagation. Parallel Distributed Processing, volume I (1986), 318-362.\n21. L. Almeida. A learning rule for asynchronous perceptrons with feedback in a combinatorial environment. IEEE First Annual International Conference on Neural Networks, 1987.\n22. F. Pineda. Generalization of back propagation to recurrent networks. Memo S1A-63-87, Johns Hopkins University Applied Physics Laboratory, 1987.\n23. W. Wickelgren. Context-sensitive coding, associative memory, and serial order in (speech) behavior. Psychological Review 76 (1969), 1-15.\n24. M.C. Mozer. Early parallel processing in reading: A connectionist approach. Attention and Performance XII (1987), 83-104.\n25. M.C. Mozer. The perception of multiple objects: A parallel, distributed processing approach. ICS Technical Report 8803, University of California, San Diego, 1988.\n26. Y. Miyata. The learning and planning of actions. ICS Technical Report 8802, University of California, San Diego, 1988.\n27. J. Bachrach. Learning to represent state. Unpublished master's thesis, University of Massachusetts, Amherst, 1988.\n28. J.L. Bybee and D.I. Slobin. Rules and schemas in the development and use of the English past tense. Language 58 (1982), 265-289.\n29. Steven Nowlan. Personal communication.\n30. R.J. Williams and D. Zipser. Experimental analysis of the real-time recurrent learning algorithm. Connection Science 1, in press.\n31. M. Gori, Y. Bengio and R. De Mori. BPS: A learning algorithm for capturing the dynamic nature of speech. Proceedings of the 1989 First International Joint Conference on Neural Networks, volume 2 (1989), 417-423.\n32. Yoshiro Miyata. Personal communication.\n"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Scanned bibliography pages processed with Windows OCR. The readable citation metadata was checked against the images and OCR; raw OCR is retained with page numbers and hashes. Spelling and source-publication inconsistencies may remain. Only reviewed Atlas matches become citation links."
      ],
      "OcrPages": [
        {
          "PdfPage": 16,
          "Text": "References\n(11 D.E. Rumelhart and J.L. McClelland, \"On learning the past tenses of English\nverbs,\" in Parallel distributed processing: FÄp/orations in the microstructure\nOf cognition. Vol. II: Psychological and biological models, J.L. McClelland\nand D.E. Rumelhart, eds., (MIT Press/ Bradford Books, Cambridge, 1986)\n216-271.\n121 T.J. Sejnowski and C.R. Rcxsenberg, \"Parallel networks that learn to pro-\nnounCe English text,\" Complex Systems, I (1987) 145—168.\n[31 G. Hinton, \"Learning distributed representations Of concepts,\" Proceedings\nof the Eighth Annual Conference of the Cognitive Science Society, Hillsdale,\nNJ (1987) 1-12.\nP. Smolensky, \"Schema selection and stochastic inference in modular envi-\nronments,\" Proceedings Of the Sixth Annual Conference on Artificial Intel-\nligence AAABS3 (1983) 109-113.\n15] J. Freyd, \"Dynamic mental representations,\" Psychological Review, 94\n(1987) 427-438.\n(61 J.L. Elman and J McClelland, \"Exploiting lawful variability in the speech\nwave,\" in Invariance and variability in speech processes, J.S. Pcrkell and\nD.II. Klatt, eds. (Erlbaum Associates, Hillsdale, NJ, 1986) 360—380.",
          "Receipt": {
            "TextSha256": "4967581BDFBF2D32C7EFAF5AC4CFF90BB1101E94B9768F29484800E04607BDD3",
            "ProcessedAtUtc": "2026-09-16T22:50:32.551099Z",
            "ImageSha256": "B61ADB719F7A658D7ED1F7E1694509952D59D107A7DBBD61F6335D7EAFB274B6",
            "Language": "en-US",
            "Lines": 53,
            "Engine": "Windows.Media.Ocr",
            "Output": "focused-refs-000016.ocr.txt",
            "Image": "focused-refs-000016.png"
          }
        },
        {
          "PdfPage": 17,
          "Text": "Michael C Mozcr\n(7) J.L. Elman and D. Zipser, \"Learning the hidden structure Of speech,\" Journal\nof the Acoustical Society Of America, in press.\n[81 T.K. Landaucr, C.A. Kamm, and S. a minimally struc,\ntured back propagation network to recognize speech,\" Proreedings of file\nNinth Annual Conference of the Ccgnitive Science %ciety, Hillsdale, N.I\n(1987) 531-536.\n[9] A. Lapedes and R. Farber, \"Nonlinear signal processing using neural net-\nworks,\" Report No. LA-UR-87-2662, Los Alamos, NM (1987).\n[101 J.L. McClelland and J L, Elman, \"Interactive processes in speech perception;\nThe TRACE model,\" Parallel distributed processing: Explorations in the\nmicrostructure of cognition. Vol. II: Psychological and biological models,\nJ.L. McClelland and D.E. Rumelhart, eds. (MIT Press, Cambridge, 1986)\n58-121.\n[Ill D.C. Plaut, S. Nowlan, and Hinton, \"Experiments on learning by back\npropagation,\" Technical Report CMU, Carnegie•Mellon University, Depart-\nmont or Computer Science, Pittsburgh, PA (1986).\n1121 D. Tank and J. Ilopfield, Proceedings of the National Academy Of Sciences,\n84 (1987) 1896.\nA. Waibel, T. Ilanazawa, G. Hinton, K. Shikano, and K. Lang, \"Phoneme\nrecognition using time-delay neural networks,\" Technical Report 1-0006,\nATR Interpreting Telephony Research Labs, Japan (1987).\nG. Ilinton, \"Connectionist learning procedures,\" Artificial Intelligence\n(1988) in press.\n115} K. Lang, \"Connectionist speech recognition,\" Unpublished Ph.D. thesis pro•\npmsaJ, Carnegie, Mellon University, Pittsburgh, PA (1987).\nJ.[,_ Elman, \"Finding structure in time,\" CRL Technical Report 8801, Uni-\nversity of California, San Diego, Center for Research in Language (1988).\nW.S. Stornetta, T. Hogg, and B.A. Huberman, \"A dynamical approach to\ntemporal pattern processing,\" Proceedings Of the IEEE Conference on Neu-\nral Information Processing Systems, Denver, CO (1987).\n1181 R.L. Watrous and L. Shastri, \"Learning acoustic features from speech data\nusing connectionist networks,\" Proceedings or the Ninth Annual Conference\nof the Ccgnitive Science Society, Hillsdale, NJ (1987) 518—530.\n( 191 M.I. Jordan, \"Attractor dynamics and parallelism in a connectionist soqllon-\ntial machine,\" Proceedings of the Fighth Annual Conference of the Cognitive\nScience Society, Hillsdale, NJ (1987) 531—546.\nA Focused Backpropagation Algorithm\n1201\n[21]\n(221\n[2.31\n1251\n130]\n[32)\nD.E. Rumelhart, G.E. Hinton, and R.J. Williams, \"Learning internal rcp-\nresentations by error propagation,\" Paral!cl distributed processing: Exp;o-\nrations in the microstructure of cognition. Vol I: Foundations, D.E.\nhart and J.L. McClelland, eds, (MIT Books, Cambridge,\n1986) 318-362.\nL. Almeida, \"A learning rule for asynchronous perceptrons with feedback in\ncombinatorial environment,\" IEEE First Annual International Conference\non Neural Networks, San Diego, CA (1987).\nF. Pineda, \"Generalization of back propagation to recurrent neural net\nworks,\" Memo SIA-63•87, Johns Hopkins University, Applied Physics Lab,\noratory, Laurel. MI) (1987).\nW. Wickelgren, \"Context-sensitive coding, associative memory, and sorial\norder in (speech) behavior,\" Review, 76 (1969) I\n-15.\nM.C. Mozer, \"Early parallel processing in reading: A connectionist ap•\nproach,n Attention and performance XIE The psychology or reading,\nM, Coltheart, ed. (Erlbaum, Hillsdale, NJ, 1987) 83-10\".\nM.C. Mozer, \"Thc perception or multiple objects: A parallel, distributed\nprocessing approach,\" ICS Technical Report 8803, University or California,\nSan Diego, Institute for Cognitive Science ( 198B).\nY. Miyata, \"The learning and planning Of actions,\" ICS Technical Report\n8802, University or California, San Diego, Institute for Cognitive Science, La\nJolla (1988).\nJ. Bachrach, \"Learning to represent state,\" master's thesis,\nUniversity of Massachusetts, Amherst (1988),\nJ.L. Bybee and D.I. Slobin, \"Rules and schemas in the development and use\nof the English past tense,\" Language, 68 (1982) 265—289.\nSteven Nowlan, personal communication.\nR.J. Williams and D. Zipser, \"Experimental analysis the real-time recur,\nrent learning algorithm,\" Connection Science. I (in press).\nM. Gori, Y. Bengio, and R. Mori, \"BPS: A learning algorithm ror capturing\nthe dynamic nature of speech,\" Proceedings Of the First International Joint\nConference on Neural Networks, 2 (1989) 417-423.\nYoshiro Miyata, personal communication.",
          "Receipt": {
            "TextSha256": "759B9480AF4BEB778FED5221C5AD3F66383A787D1CFD2C8D4A06EDA06235E247",
            "ProcessedAtUtc": "2026-09-16T22:50:32.9506732Z",
            "ImageSha256": "25CC09B1861F01BD809B33B40FB050AF19608C51F185F0D5A9887515CBF20BC0",
            "Language": "en-US",
            "Lines": 79,
            "Engine": "Windows.Media.Ocr",
            "Output": "focused-refs-000017.ocr.txt",
            "Image": "focused-refs-000017.png"
          }
        }
      ]
    },
    {
      "Slug": "real-time-recurrent-sequences",
      "Paper": "Learning Sequential Structure with the Real-Time Recurrent Learning Algorithm",
      "AtlasYear": 1989,
      "Status": "unavailable",
      "Method": "unavailable",
      "SourceUrl": "https://doi.org/10.1142/S0129065789000037",
      "PdfSha256": null,
      "Sections": [],
      "PublisherReferences": [],
      "Notes": [
        "The publisher record supplies no reference list through Crossref. After human verification, the publisher page was inspected: its References tab says None and the article is marked No Access. The abstract contains numbered citations, so the empty HTML reference tab does not mean no bibliography. Alternate repository and author-copy searches have not recovered the seven-page paper. The original reference page is still required."
      ]
    },
    {
      "Slug": "the-ai-toy",
      "Paper": "Future Computing: Neural Networks, Parts 1, 2, and 3",
      "AtlasYear": 1990,
      "Status": "indexed",
      "Method": "windows-ocr-with-visual-source-box-review",
      "SourceUrl": "https://archive.org/details/1990-03-computegazette/page/n45/mode/1up",
      "PdfSha256": "B69EF94A0F859EB5E55F2D84FD501A8FFD728A8E81DA97AEC976403E50E3B96C",
      "Sections": [
        {
          "Section": "Sources, Part 3, printed page 44",
          "PdfPages": [
            46
          ],
          "Text": "Neural Computing: Theory and Practice. By Philip Wasserman. Van Nostrand Reinhold.\nNeurocomputing: Foundations of Research. Edited by James A. Anderson and Edward Rosenfeld. MIT Press.\nParallel Distributed Processing (two volumes); Explorations in Parallel Distributed Processing. By Rumelhart, McClelland, and the PDP Research Group. MIT Press. [These titles appear together in one source block in the magazine.]"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "The March 1990 installment contains a Sources box, which was previously missed. All three printed source blocks are transcribed and checked against the rendered page after OCR. The final block names Parallel Distributed Processing and Explorations in Parallel Distributed Processing together. These books are distinct from the 1986 Nature backpropagation paper and do not create a link to that entry."
      ],
      "BibliographyEdition": "March 1990, Part 3 of the three-part magazine series",
      "OcrPages": [
        {
          "PdfPage": 46,
          "Text": "Neural Computing: Theory and Practice. By Philip Wasserman. Van Nostrand Reinhold.\nNeurocomputing: Foundations of Research. Edited by James A. Anderson and Edward Rosenfeld. MIT Press.\nParallel Distributed Processing (two volumes); Explorations in Parallel Distributed Processing. By Rumelhart, McClelland, and the PDP Research Group. MIT Press. [These titles appear together in one source block in the magazine.]",
          "TextScope": "Reviewed Sources box transcription only; receipt hashes identify the retained full-page OCR artifact.",
          "Receipt": {
            "TextSha256": "004DE883D603EDF84C387AC19EDF4BDBB239B61E1EBC974C510FC9A8772B0FD0",
            "ProcessedAtUtc": "2026-09-16T23:55:19.0280491Z",
            "ImageSha256": "AD26C236B863AD68A5FE5DEF90D26128D539953B2012B2554484769D1D91DACE",
            "Language": "en-US",
            "Lines": 436,
            "Engine": "Windows.Media.Ocr",
            "Output": "ai-toy-refs-000046.ocr.txt",
            "Image": "ai-toy-refs-000046.png"
          }
        }
      ],
      "ReferenceBlockCount": 3
    },
    {
      "Slug": "support-vector-networks",
      "Paper": "Support-vector networks",
      "AtlasYear": 1995,
      "Status": "indexed",
      "Method": "publisher-reference-section",
      "SourceUrl": "https://link.springer.com/article/10.1007/BF00994018",
      "PdfSha256": null,
      "Sections": [
        {
          "Section": "Publisher references",
          "StartLine": 1,
          "EndLine": 14,
          "PdfPages": [],
          "Text": "Aizerman, M., Braverman, E., & Rozonoer, L. (1964). Theoretical foundations of the potential function method in pattern recognition learning.Automation and Remote Control, 25:821–837.\nAnderson, T.W., & Bahadur, R.R. (1966). Classification into two multivariate normal distributions with different covariance matrices.Ann. Math. Stat., 33:420–431.\nBoser, B.E., Guyon, I., & Vapnik, V.N. (1992). A training algorithm for optimal margin classifiers. InProceedings of the Fifth Annual Workshop of Computational Learning Theory, 5, 144–152. Pittsburgh, ACM.\nBottou, L., Cortes, C., Denker, J.S., Drucker, H., Guyon, I., Jackel, L.D., LeCun, Y., Sackinger, E., Simard, P., Vapnik, V., & Miller, U.A. (1994). Comparison of classifier methods: A case study in handwritten digit recognition.Proceedings of 12th International Conference on Pattern Recognition and Neural Network.\nBromley, J., & Sackinger, E. (1991). Neural-network andk-nearest-neighbor classifiers. Technical Report 11359-910819-16TM, AT&T.\nCournant, R., & Hilbert, D. (1953).Methods of Mathematical Physics, Interscience, New York.\nFisher, R.A. (1936). The use of multiple measurements in taxonomic problems.Ann. Eugenics, 7:111–132.\nLeCun, Y. (1985). Une procedure d'apprentissage pour reseau a seuil assymetrique.Cognitiva 85: A la Frontiere de l'Intelligence Artificielle des Sciences de la Connaissance des Neurosciences, 599–604, Paris.\nLeCun, Y., Boser, B., Denker, J.S., Henderson, D., Howard, R.E., Hubbard, W., & Jackel, L.D. (1990). Handwritten digit recognition with a back-propagation network.Advances in Neural Information Processing Systems, 2, 396–404, Morgan Kaufman.\nParker, D.B. (1985). Learning logic. Technical Report TR-47, Center for Computational Research in Economics and Management Science, Massachusetts Institute of Technology, Cambridge, MA.\nRosenblatt, F. (1962).Principles of Neurodynamics, Spartan Books, New York.\nRumelhart, D.E., Hinton, G.E., & Williams, R.J. (1986). Learning internal representations by backpropagating errors.Nature, 323:533–536.\nRumelhart, D.E., Hinton, G.E., & Williams, R.J. (1987). Learning internal representations by error propagation. In James L. McClelland & David E. Rumelhart (Eds.),Parallel Distributed Processing, 1, 318–362, MIT Press.\nVapnik, V.N. (1982).Estimation of Dependences Based on Empirical Data, Addendum 1, New York: Springer-Verlag."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Reference section extracted from the linked source. Only reviewed matches to existing Atlas works become citation links; extraction may retain typographic or column-order artifacts."
      ]
    },
    {
      "Slug": "lstm",
      "Paper": "Long Short-Term Memory",
      "AtlasYear": 1997,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.bioinf.jku.at/publications/older/2604.pdf",
      "PdfSha256": "CEB9E53DBC0493F5B3BF5520ED940F3E6B526064D17B2118D77E51F79C0EDCC6",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 2142,
          "EndLine": 2186,
          "PdfPages": [
            30,
            31,
            32
          ],
          "Text": "Almeida, L. B. (1987). A learning rule for asynchronous perceptrons with feedback in a combinatorial environment. In IEEE 1st International Conference on Neural Networks, San Diego, volume 2, pages 609{618.\nBaldi, P. and Pineda, F. (1991). Contrastive learning and neural oscillator. Neural Computation, 3:526{545.\nBengio, Y. and Frasconi, P. (1994). Credit assignment through time: Alternatives to backpropagation. In Cowan, J. D., Tesauro, G., and Alspector, J., editors, Advances in Neural Information Processing Systems 6, pages 75{82. San Mateo, CA: Morgan Kaufmann.\nBengio, Y., Simard, P., and Frasconi, P. (1994). Learning long-term dependencies with gradient descent is di cult. IEEE Transactions on Neural Networks, 5(2):157{166.\nCleeremans, A., Servan-Schreiber, D., and McClelland, J. L. (1989). Finite-state automata and simple recurrent networks. Neural Computation, 1:372{381.\nde Vries, B. and Principe, J. C. (1991). A theory for neural networks with time delays. In Lippmann, R. P., Moody, J. E., and Touretzky, D. S., editors, Advances in Neural Information Processing Systems 3, pages 162{168. San Mateo, CA: Morgan Kaufmann.\nDoya, K. (1992). Bifurcations in the learning of recurrent neural networks. In Proceedings of 1992 IEEE International Symposium on Circuits and Systems, pages 2777{2780.\nDoya, K. and Yoshizawa, S. (1989). Adaptive neural oscillator using continuous-time backpropagation learning. Neural Networks, 2:375{385.\nElman, J. L. (1988). Finding structure in time. Technical Report CRL Technical Report 8801, Center for Research in Language, University of California, San Diego.\nFahlman, S. E. (1991). The recurrent cascade-correlation learning algorithm. In Lippmann, R. P., Moody, J. E., and Touretzky, D. S., editors, Advances in Neural Information Processing Systems 3, pages 190{196. San Mateo, CA: Morgan Kaufmann.\nHochreiter, J. (1991). Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, Institut fur Informatik, Lehrstuhl Prof. Brauer, Technische Universitat Munchen. See www7.informatik.tu-muenchen.de/~hochreit.\n30\n\n\fHochreiter, S. and Schmidhuber, J. (1995). Long short-term memory. Technical Report FKI-20795, Fakultat fur Informatik, Technische Universitat Munchen.\nHochreiter, S. and Schmidhuber, J. (1996). Bridging long time lags by weight guessing and \\Long Short-Term Memory\". In Silva, F. L., Principe, J. C., and Almeida, L. B., editors, Spatiotemporal models in biological and arti cial systems, pages 65{72. IOS Press, Amsterdam, Netherlands. Serie: Frontiers in Arti cial Intelligence and Applications, Volume 37.\nHochreiter, S. and Schmidhuber, J. (1997). LSTM can solve hard long time lag problems. In Advances in Neural Information Processing Systems 9. MIT Press, Cambridge MA. Presented at NIPS 96.\nLang, K., Waibel, A., and Hinton, G. E. (1990). A time-delay neural network architecture for isolated word recognition. Neural Networks, 3:23{43.\nMiller, C. B. and Giles, C. L. (1993). Experimental comparison of the e ect of order in recurrent neural networks. International Journal of Pattern Recognition and Arti cial Intelligence, 7(4):849{872.\nMozer, M. C. (1989). A focused back-propagation algorithm for temporal sequence recognition. Complex Systems, 3:349{381.\nMozer, M. C. (1992). Induction of multiscale temporal structure. In Lippman, D. S., Moody, J. E., and Touretzky, D. S., editors, Advances in Neural Information Processing Systems 4, pages 275{282. San Mateo, CA: Morgan Kaufmann.\nPearlmutter, B. A. (1989). Learning state space trajectories in recurrent neural networks. Neural Computation, 1(2):263{269.\nPearlmutter, B. A. (1995). Gradient calculations for dynamic recurrent neural networks: A survey. IEEE Transactions on Neural Networks, 6(5):1212{1228.\nPineda, F. J. (1987). Generalization of back-propagation to recurrent neural networks. Physical Review Letters, 19(59):2229{2232.\nPineda, F. J. (1988). Dynamics and architecture for neural computation. Journal of Complexity, 4:216{245.\nPlate, T. A. (1993). Holographic recurrent networks. In S. J. Hanson, J. D. C. and Giles, C. L., editors, Advances in Neural Information Processing Systems 5, pages 34{41. San Mateo, CA: Morgan Kaufmann.\nPollack, J. B. (1991). Language induction by phase transition in dynamical recognizers. In Lippmann, R. P., Moody, J. E., and Touretzky, D. S., editors, Advances in Neural Information Processing Systems 3, pages 619{626. San Mateo, CA: Morgan Kaufmann.\nPuskorius, G. V. and Feldkamp, L. A. (1994). Neurocontrol of nonlinear dynamical systems with Kalman lter trained recurrent networks. IEEE Transactions on Neural Networks, 5(2):279{ 297.\nRing, M. B. (1993). Learning sequential tasks by incrementally adding higher orders. In S. J. Hanson, J. D. C. and Giles, C. L., editors, Advances in Neural Information Processing Systems 5, pages 115{122. Morgan Kaufmann.\nRobinson, A. J. and Fallside, F. (1987). The utility driven dynamic error propagation network. Technical Report CUED/F-INFENG/TR.1, Cambridge University Engineering Department.\nSchmidhuber, J. (1989). The Neural Bucket Brigade: A local learning algorithm for dynamic feedforward and recurrent networks. Connection Science, 1(4):403{412.\n31\n\n\fSchmidhuber, J. (1992a). A xed size storage O(n3) time complexity learning algorithm for fully recurrent continually running networks. Neural Computation, 4(2):243{248.\nSchmidhuber, J. (1992b). Learning complex, extended sequences using the principle of history compression. Neural Computation, 4(2):234{242.\nSchmidhuber, J. (1992c). Learning unambiguous reduced sequence descriptions. In Moody, J. E., Hanson, S. J., and Lippman, R. P., editors, Advances in Neural Information Processing Systems 4, pages 291{298. San Mateo, CA: Morgan Kaufmann.\nSchmidhuber, J. (1993). Netzwerkarchitekturen, Zielfunktionen und Kettenregel. Habilitationsschrift, Institut fur Informatik, Technische Universitat Munchen.\nSchmidhuber, J. and Hochreiter, S. (1996). Guessing can outperform many long time lag algorithms. Technical Report IDSIA-19-96, IDSIA.\nSilva, G. X., Amaral, J. D., Langlois, T., and Almeida, L. B. (1996). Faster training of recurrent networks. In Silva, F. L., Principe, J. C., and Almeida, L. B., editors, Spatiotemporal models in biological and arti cial systems, pages 168{175. IOS Press, Amsterdam, Netherlands. Serie: Frontiers in Arti cial Intelligence and Applications, Volume 37.\nSmith, A. W. and Zipser, D. (1989). Learning sequential structures with the real-time recurrent learning algorithm. International Journal of Neural Systems, 1(2):125{131.\nSun, G., Chen, H., and Lee, Y. (1993). Time warping invariant neural networks. In S. J. Hanson, J. D. C. and Giles, C. L., editors, Advances in Neural Information Processing Systems 5, pages 180{187. San Mateo, CA: Morgan Kaufmann.\nWatrous, R. L. and Kuhn, G. M. (1992). Induction of nite-state languages using second-order recurrent networks. Neural Computation, 4:406{414.\nWerbos, P. J. (1988). Generalization of backpropagation with application to a recurrent gas market model. Neural Networks, 1.\nWilliams, R. J. (1989). Complexity of exact gradient computation algorithms for recurrent neural networks. Technical Report Technical Report NU-CCS-89-27, Boston: Northeastern University, College of Computer Science.\nWilliams, R. J. and Peng, J. (1990). An e cient gradient-based algorithm for on-line training of recurrent network trajectories. Neural Computation, 4:491{501.\nWilliams, R. J. and Zipser, D. (1992). Gradient-based learning algorithms for recurrent networks and their computational complexity. In Back-propagation: Theory, Architectures and Applications. Hillsdale, NJ: Erlbaum."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "document-recognition",
      "Paper": "Gradient-Based Learning Applied to Document Recognition",
      "AtlasYear": 1998,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://raw.githubusercontent.com/tpn/pdfs/master/Gradient-Based%20Learning%20Applied%20to%20Document%20Recognition%20-%201998%20%28Lecun98%29.pdf",
      "PdfSha256": "3C2CEB4CF8F44FD7C489BA9712DA9C751B7623493DEF2312274DF7CB3CED2688",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 2671,
          "EndLine": 3040,
          "PdfPages": [
            43,
            44,
            45,
            46
          ],
          "Text": "[1] R. O. Duda and P. E. Hart, Pattern Classiﬁcation and Scene Analysis. New York: Wiley, 1973.\n[2] Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural Computation, vol. 1, no. 4, pp. 541–551, Winter 1989.\n[3] S. Seung, H. Sompolinsky, and N. Tishby, “Statistical mechanics of learning from examples,” Phys. Rev. A, vol. 45, pp. 6056–6091, 1992.\n[4] V. N. Vapnik, E. Levin, and Y. LeCun, “Measuring the vcdimension of a learning machine,” Neural Computation, vol. 6, no. 5, pp. 851–876, 1994.\n[5] C. Cortes, L. Jackel, S. Solla, V. N. Vapnik, and J. Denker, “Learning curves: Asymptotic values and rate of convergence,”\nPROCEEDINGS OF THE IEEE, VOL. 86, NO. 11, NOVEMBER 1998\n\n\fin Advances in Neural Information Processing Systems 6, J. D.\n\nCowan, G. Tesauro, and J. Alspector, Eds. San Mateo, CA:\n\nMorgan Kaufmann, 1994, pp. 327–334.\n\n[6] V. N. Vapnik, The Nature of Statistical Learning Theory. New\n\nYork: Springer, 1995.\n\n[7]\n\n, Statistical Learning Theory. New York: Wiley, 1998.\n\n[8] W. H. Press, B. P. Flannery, S. A. Teukolsky, and W. T.\n\nVetterling, Numerical Recipes: The Art of Scientiﬁc Computing.\n\nCambridge, UK: Cambridge Univ., 1986.\n\n[9] S. I. Amari, “A theory of adaptive pattern classiﬁers,” IEEE\n\nTrans. Electron. Comput., vol. EC-16, pp. 299–307, 1967.\n\n[10] Y. Tsypkin, Adaptation and Learning in Automatic Systems\n\nNew York: Academic, 1971.\n\n[11]\n\n, Foundations of the Theory of Learning Systems. New\n\nYork: Academic, 1973.\n\n[12] M. Minsky and O. Selfridge, “Learning in random nets,” in\n\nProc. 4th London Symp. Information Theory, pp. 335–347, 1961.\n\n[13] D. H. Ackley, G. E. Hinton, and T. J. Sejnowski, “A learning\n\nalgorithm for Boltzmann machines,” Cognitive Sci., vol. 9, pp.\n\n147–169, 1985.\n\n[14] G. E. Hinton and T. J. Sejnowski, “Learning and relearning\n\nin Boltzmann machines,” in Parallel Distributed Processing:\n\nExplorations in the Microstructure of Cognition. Volume 1:\n\nFoundations, D. E. Rumelhart and J. L. McClelland, Eds.\n\nCambridge, MA: MIT, 1986.\n\n[15] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learn-\n\ning internal representations by error propagation,” in Parallel\n\nDistributed Processing: Explorations in the Microstructure of\n\nCognition, vol. I. Cambridge, MA: Bradford Books, 1986,\n\npp. 318–362,\n\n[16] A. E. Bryson, Jr. and Y.-C. Ho, Applied Optimal Control.\n\nLondon, UK: Blaisdell, 1969.\n\n[17] Y. LeCun, “A learning scheme for asymmetric threshold\n\nnetworks,” in Proc. Cognitiva ’85, Paris, France, 1985, pp.\n\n599–604.\n\n[18]\n\n, “Learning processes in an asymmetric threshold net-\n\nwork,” in Disordered Systems and Biological Organization, E.\n\nBienenstock, F. Fogelman-Soulie¨, and G. Weisbuch, Eds. Les\n\nHouches, France: Springer-Verlag, 1986, pp. 233–240.\n\n[19] D. B. Parker, “Learning-logic,” Sloan School Manage., MIT,\n\nCambridge, MA, Tech. Rep., TR-47, Apr. 1985.\n\n[20] Y. LeCun, Mode´les Connexionnistes de l’Apprentissage (Con-\n\nnectionist Learning Models), Ph.D. dissertation, Universite´ P.\n\net M. Curie (Paris 6), June 1987.\n\n[21]\n\n, “A theoretical framework for back-propagation,” in Proc.\n\n1988 Connectionist Models Summer School, D. Touretzky, G.\n\nHinton, and T. Sejnowski, Eds. Pittsburgh, PA: CMU, Morgan\n\nKaufmann, 1988, pp. 21–28.\n\n[22] L. Bottou and P. Gallinari, “A framework for the cooperation\n\nof learning algorithms,” in Advances in Neural Information\n\nProcessing Systems, vol. 3, D. Touretzky and R. Lippmann,\n\nEds. Denver, CO: Morgan Kaufmann, 1991.\n\n[23] C. Y. Suen, C. Nadal, R. Legault, T. A. Mai, and L. Lam,\n\n“Computer recognition of unconstrained handwritten numerals,”\n\nProc. IEEE, vol. 80, pp. 1162–1180, July 1992.\n\n[24] S. N. Srihari, “High-performance reading machines,” Proc.\n\nIEEE., vol. 80, pp. 1120–1132, July 1992.\n\n[25] Y. LeCun, L. D. Jackel, B. Boser, J. S. Denker, H. P. Graf,\n\nI. Guyon, D. Henderson, R. E. Howard, and W. Hubbard,\n\n“Handwritten digit recognition: Applications of neural net chips\n\nand automatic learning,” IEEE Trans. Commun., vol. 37, pp.\n\n41–46, Nov. 1989.\n\n[26] J. Keeler, D. Rumelhart, and W. K. Leow, “Integrated seg-\n\nmentation and recognition of hand-printed numerals,” in Neural\n\nInformation Processing Systems, R. P. Lippmann, J. M. Moody,\n\nand D. S. Touretzky, Eds. San Mateo, CA: Morgan Kaufmann,\n\nvol. 3, pp. 557–563, 1991.\n\n[27] O. Matan, C. J. C. Burges, Y. LeCun, and J. S. Denker, “Multi-\n\ndigit recognition using a space displacement neural network,”\n\nvol. 4, in Neural Information Processing Systems, J. M. Moody,\n\nS. J. Hanson, and R. P. Lippman, Eds. San Mateo, CA:\n\nMorgan Kaufmann, 1992.\n\n[28] L. R. Rabiner, “A tutorial on hidden Markov models and\n\nselected applications in speech recognition,” Proc. IEEE, vol.\n\n77, pp. 257–286, Feb. 1989.\n\n[29] H. A. Bourland and N. Morgan, Connectionist Speech Recog-\n\nnition: A Hybrid Approach. Boston: Kluwer, 1994.\n\n[30] D. H. Hubel and T. N. Wiesel, “Receptive ﬁelds, binocular in-\n\nteraction, and functional architecture in the cat’s visual cortex,”\n\nJ. Physiology (London), vol. 160, pp. 106–154, 1962.\n\n[31] K. Fukushima, “Cognition: A self-organizing multilayered neu-\n\nral network,” Biological Cybern., vol. 20, pp. 121–136, 1975.\n\n[32] K. Fukushima and S. Miyake, “Neocognitron: A new algorithm\n\nfor pattern recognition tolerant of deformations and shifts in\n\nposition,” Pattern Recognit., vol. 15, no. 6, pp. 455–469, Nov.\n\n1982.\n\n[33] M. C. Mozer, The Perception of Multiple Objects: A Con-\n\nnectionist Approach. Cambridge, MA: MIT-Bradford Books,\n\n1991.\n\n[34] Y. LeCun, “Generalization and network design strategies,”\n\nin Connectionism in Perspective, R. Pfeifer, Z. Schreter, F.\n\nFogelman, and L. Steels, Eds. Zurich, Switzerland: Elsevier,\n\n1989.\n\n[35] Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E.\n\nHoward, W. Hubbard, and L. D. Jackel, “Handwritten digit\n\nrecognition with a back-propagation network,” in Advances\n\nin Neural Information Processing Systems 2 (NIPS’89), David\n\nTouretzky, Ed. Denver, CO: Morgan Kaufmann, 1990.\n\n[36] G. L. Martin, “Centered-object integrated segmentation and\n\nrecognition of overlapping hand-printed characters,” Neural\n\nComputation, vol. 5, no. 3, pp. 419–429, 1993.\n\n[37] J. Wang and J. Jean, “Multi-resolution neural networks for\n\nomnifont character recognition,” in Proc. Int. Conf. Neural\n\nNetworks, vol. III, 1993, pp. 1588–1593.\n\n[38] Y. Bengio, Y. LeCun, C. Nohl, and C. Burges, “Lerec: A\n\nNN/HMM hybrid for on-line handwriting recognition,” Neural\n\nComputation, vol. 7, no. 5, 1995.\n\n[39] S. Lawrence, C. L. Giles, A. C. Tsoi, and A. D. Back, “Face\n\nrecognition: A convolutional neural network approach,” IEEE\n\nTrans. Neural Networks, vol. 8, pp. 98–113, Jan. 1997.\n\n[40] K. J. Lang and G. E. Hinton, “A time delay neural network\n\narchitecture for speech recognition,” Carnegie-Mellon Univ.,\n\nPittsburgh, PA, Tech. Rep. CMU-CS-88-152, 1988.\n\n[41] A. H. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K.\n\nLang, “Phoneme recognition using time-delay neural networks,”\n\nIEEE Trans. Acoustics, Speech, Signal Processing, vol. 37, pp.\n\n328–339, Mar. 1989.\n\n[42] L. Bottou, F. Fogelman, P. Blanchet, and J. S. Lienard, “Speaker\n\nindependent isolated digit recognition: Multilayer perceptron\n\nversus dynamic time warping,” Neural Networks, vol. 3, pp.\n\n453–465, 1990.\n\n[43] P. Haffner and A. H. Waibel, “Time-delay neural networks\n\nembedding time alignment: A performance analysis,” in Proc.\n\nEUROSPEECH’91, 2nd Europ. Conf. Speech Communication\n\nand Technology, Genova, Italy.\n\n[44] I. Guyon, P. Albrecht, Y. LeCun, J. S. Denker, and W. Hubbard,\n\n“Design of a neural network character recognizer for a touch\n\nterminal,” Pattern Recognit., vol. 24, no. 2, pp. 105–119, 1991.\n\n[45] J. Bromley, J. W. Bentz, L. bottou, I. Guyon, Y. LeCun, C.\n\nMoore, E. Sa¨ckinger, and R. Shah, “Signature veriﬁcation using\n\na siamese time delay neural network,” Int. J. Pattern Recognit.\n\nArtiﬁcial Intell., vol. 7, no. 4, pp. 669–687, Aug. 1993.\n\n[46] Y. LeCun, I. Kanter, and S. Solla, “Eigenvalues of covariance\n\nmatrices: Application to neural-network learning,” Phys. Rev.\n\nLett., vol. 66, no. 18, pp. 2396–2399, May 1991.\n\n[47] T. G. Dietterich and G. Bakiri, “Solving multiclass learning\n\nproblems via error-correcting output codes,” J. Artiﬁcial Intell.\n\nRes., vol. 2, pp. 263–286, 1995.\n\n[48] L. R. Bahl, P. F. Brown, P. V. de Souza, and R. L. Mercer,\n\n“Maximum mutual information of hidden Markov model pa-\n\nrameters for speech recognition,” in Proc. Int. Conf. Acoustics,\n\nSpeech, Signal Processing, 1986, pp. 49–52.\n\n[49]\n\n, “Speech recognition with continuous-parameter hidden\n\nMarkov models,” Comput., Speech Language, vol. 2, pp.\n\n219–234, 1987.\n\n[50] B. H. Juang and S. Katagiri, “Discriminative learning for\n\nminimum error classiﬁcation,” IEEE Trans. Acoustics, Speech,\n\nSignal Processing, vol. 40, pp. 3043–3054, Dec. 1992.\n\n[51] Y. LeCun, L. D. Jackel, L. Bottou, A. Brunot, C. Cortes, J. S.\n\nDenker, H. Drucker, I. Guyon, U. A. Muller, E. Sa¨ckinger, P.\n\nSimard, and V. N. Vapnik, “Comparison of learning algorithms\n\nfor handwritten digit recognition,” in Int. Conf. Artiﬁcial Neural\n\nNetworks, F. Fogelman and P. Gallinari, Eds. Paris: EC2 &\n\nCie, 1995, pp. 53–60.\n\n[52] I. Guyon, I. Poujaud, L. Personnaz, G. Dreyfus, J. Denker, and\n\nY. LeCun, “Comparing different neural net architectures for\n\nLECUN et al.: GRADIENT-BASED LEARNING APPLIED TO DOCUMENT RECOGNITION\n\n2321\n\n\fclassifying handwritten digits,” in Proc. IEEE IJCNN, Washington, DC, vol. II, 1989, pp. 127–132,. [53] R. Ott, “Construction of quadratic polynomial classiﬁers,” in Proc. IEEE Int. Conf. Pattern Recognition, 1976, pp. 161–165. [54] J. Schu¨rmann, “A multifont word recognition system for postal address reading,” IEEE Trans. Comput., vol. C-27, pp. 721–732, Aug. 1978. [55] Y. Lee, “Handwritten digit recognition using k-nearest neighbor, radial-basis functions, and backpropagation neural networks,” Neural Computation, vol. 3, no. 3, pp. 440–449, 1991. [56] D. Saad and S. A. Solla, “Dynamics of on-line gradient descent learning for multilayer neural networks,” in Advances in Neural Information Processing Systems, vol. 8, D. S. Touretzky, M. C. Mozer, and M. E. Hasselmo, Eds. Cambridge, MA: MIT, 1996, pp. 302–308. [57] G. Cybenko, “Approximation by superpositions of sigmoidal functions,” Math. Control, Signals, Syst., vol. 2, no. 4, pp. 303–314, 1989. [58] L. Bottou and V. N. Vapnik, “Local learning algorithms,” Neural Computation, vol. 4, no. 6, pp. 888–900, 1992. [59] R. E. Schapire, “The strength of weak learnability,” Machine Learning, vol. 5, no. 2, pp. 197–227, 1990. [60] H. Drucker, R. Schapire, and P. Simard, “Improving performance inneural networks using a boosting algorithm,” in Advances in Neural Information Processing Systems 5, S. J. Hanson, J. D. Cowan, and C. L. Giles, Eds. San Mateo, CA: Morgan Kaufmann, 1993, pp. 42–49. [61] P. Simard, Y. LeCun, and J. Denker, “Efﬁcient pattern recognition using a new transformation distance,” in Advances in Neural Information Processing Systems, vol. 5, S. Hanson, J. Cowan, and L. Giles, Eds. San Mateo, CA: Morgan Kaufmann, 1993. [62] B. Boser, I. Guyon, and V. Vapnik, “A training algorithm for optimal margin classiﬁers,” in Proc. 5th Annu. Workshop Computational Learning Theory, vol. 5, 1992, pp. 144–152. [63] C. J. C. Burges and B. Schoelkopf, “Improving the accuracy and speed of support vector machines,” in Advances in Neural Information Processing Systems 9, M. Jordan, M. Mozer, and T. Petsche, Eds. Cambridge, MA: MIT, 1997. [64] E. Sa¨ckinger, B. Boser, J. Bromley, Y. LeCun, and L. D. Jackel, “Application of the ANNA neural network chip to high-speed character recognition,” IEEE Trans. Neural Networks, vol. 3, no. 3, pp. 498–505, Mar. 1992. [65] J. S. Bridle, “Probabilistic interpretation of feedforward classiﬁcation networks outputs, with relationship to statistical pattern recognition,” in Neurocomputing, Algorithms, Architectures and Applications, F. Fogelman, J. Herault, and Y. Burnod, Eds. Les Arcs, France: Springer, 1989. [66] Y. LeCun, L. Bottou, and Y. Bengio, “Reading checks with graph transformer networks,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Processing. Munich, Germany, vol. 1, 1997, pp. 151–154,. [67] Y. Bengio, Neural Networks for Speech and Sequence Recognition. London, UK: International Thompson, 1996. [68] C. Burges, O. Matan, Y. LeCun, J. Denker, L. Jackel, C. Stenard, C. Nohl, and J. Ben, “Shortest path segmentation: A method for training a neural network to recognize character strings,” in Proc. Int. Joint Conf. Neural Networks, Baltimore, MD, vol. 3, 1992, pp. 165–172. [69] T. M. Breuel, “A system for the off-line recognition of handwritten text,” in Proc. IEEE ICPR’94, Jerusalem, pp. 129–134. [70] A. Viterbi, “Error bounds for convolutional codes and an asymptotically optimum decoding algorithm,” IEEE Trans. Inform. Theory, vol. 15, pp. 260–269, Apr. 1967. [71] R. P. Lippmann and B. Gold, “Neural-net classiﬁers useful for speech recognition,” in Proc. IEEE 1st Int. Conf. Neural Networks, San Diego, CA, June 1987, pp. 417–422. [72] H. Sakoe, R. Isotani, K. Yoshida, K. Iso, and T. Watanabe, “Speaker-independent word recognition using dynamic programming neural networks,” in Proc. Int. Conf. Acoustics, Speech, Signal Processing, Glasgow, 1989, pp. 29–32. [73] J. S. Bridle, “Alphanets: A recurrent ‘neural’ network architecture with a hidden Markov model interpretation,” Speech Commun., vol. 9, no. 1, pp. 83–92, 1990. [74] M. A. Franzini, K. F. Lee, and A. H. Waibel, “Connectionist viterbi training: A new hybrid method for continuous speech recognition,” in Proc. Int. Conf. Acoustics, Speech, Signal Processing, Albuquerque, NM, 1990, pp. 425–428.\n2322\n\n[75] L. T. Niles and H. F. Silverman, “Combining hidden Markov models and neural network classiﬁers,” in Proc. Int. Conf. Acoustics, Speech, Signal Processing, Albuquerque, NM, 1990, pp. 417–420.\n[76] X. Driancourt and L. Bottou, “MLP, LVQ and DP: Comparison & cooperation,” in Proc. Int. Joint Conf. Neural Networks, Seattle, WA, vol. 2, 1991, pp. 815–819.\n[77] Y. Bengio, R. De Mori, G. Flammia, and R. Kompe, “Global optimization of a neural network-hidden Markov model hybrid,” IEEE Trans. Neural Networks, vol. 3, pp. 252–259, March 1992.\n[78] P. Haffner and A. H. Waibel, “Multi-state time-delay neural networks for continuous speech recognition,” vol. 4, in Advances in Neural Information Processing Systems. San Mateo, CA: Morgan Kaufmann, pp. 579–588, 1992.\n[79] Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difﬁcult,” IEEE Trans. Neural Networks, vol. 5, no. 2, pp. 157–166, Mar. 1994.\n[80] T. Kohonen, G. Barna, and R. Chrisley, “Statistical pattern recognition with neural network: Benchmarking studies,” in Proc. IEEE 2nd Int. Conf. Neural Networks, San Diego, CA, vol. 1, 1988, pp. 61–68.\n[81] P. Haffner, “Connectionist speech recognition with a global MMI algorithm,” in Proc. EUROSPEECH’93, 3rd Europ. Conf. Speech Communication and Technology, Berlin, pp. 1929–1932.\n[82] J. S. Denker and C. J. Burges, “Image segmentation and recognition,” in The Mathematics of Induction. Reading, MA: Addison Wesley, 1995.\n[83] L. Bottou, Une Approche the´orique de l’Apprentissage Connexionniste: Applications a` la Reconnaissance de la Parole, Ph.D. dissertation, Univ. Paris XI, France, 1991.\n[84] M. Rahim, Y. Bengio, and Y. LeCun, “Disriminative feature and model design for automatic speech recognition,” in Proc. Eurospeech, Rhodes, Greece, 1997, pp. 75–78.\n[85] U. Bodenhausen, S. Manke, and A. Waibel, “Connectionist architectural learning for high performance character and speech recognition,” in Proc. Int. Conf. Acoustics, Speech, Signal Processing, Minneapolis, MN, vol. 1, 1993, pp. 625–628.\n[86] F. Pereira, M. Riley, and R. Sproat, “Weighted rational transductions and their application to human language processing,” in ARPA Natural Language Processing Workshop, 1994.\n[87] M. Lades, J. C. Vorbru¨ggen, J. Buhmann, and C. von der Malsburg, “Distortion invariant object recognition in the dynamic link architecture,” IEEE Trans. Comput., vol. 42, pp. 300–311, March 1993.\n[88] B. Boser, E. Sa¨ckinger, J. Bromley, Y. LeCun, and L. Jackel, “An analog neural network processor with programmable topology,” IEEE J. Solid-State Circuits, vol. 26, pp. 2017–2025, Dec. 1991.\n[89] M. Schenkel, H. Weissman, I. Guyon, C. Nohl, and D. Henderson, “Recognition-based segmentation of on-line hand-printed words,” in Advances in Neural Information Processing Systems 5, S. J. Hanson, J. D. Cowan, and C. L. Giles, Eds. Denver, CO: Morgan Kaufmann, 1993, pp. 723–730.\n[90] C. Dugust, L. Devillers, and X. Aubert, “Combining TDNN and HMM in a hybrid system for improved continuous-speech recognition,” IEEE Trans. Speech Audio Processing, vol. 2, pp. 217–224, Jan. 1994.\n[91] O. Matan, H. S. Baird, J. Bromley, C. J. C. Burges, J. S. Denker, L. D. Jackel, Y. LeCun, E. P. D. Pednault, W. Satterﬁeld, C. E. Stenard, and T. J. Thompson, “Reading handwritten digits: A ZIP code recognition system,” IEEE Trans. Comput., vol. 25, no. 7, pp. 59–63, July 1992.\n[92] Y. Bengio and Y. LeCun, “Word normalization for on-line handwritten word recognition,” in Proc. IEEE Int. Conf. Pattern Recognition, Jerusalem, 1994.\n[93] R. Vaillant, C. Monrocq, and Y. LeCun, “Original approach for the localization of objects in images,” Proc. Inst. Elect. Eng., vol. 141, no. 4, pp. 245–250, Aug. 1994.\n[94] R. Wolf and J. Platt, “Postal address block location using a convolutional locator network,” in Advances in Neural Information Processing Systems 6, J. D. Cowan, G. Tesauro, and J. Alspector, Eds. San Mateo, CA: Morgan Kaufmann, 1994, pp. 745–752.\n[95] S. Nowlan and J. Platt, “A convolutional neural network hand tracker,” in Advances in Neural Information Processing Systems 7, G. Tesauro, D. Touretzky, and T. Leen, Eds. San Mateo, CA: Morgan Kaufmann, 1995, pp. 901–908.\n[96] H. A. Rowley, S. Baluja, and T. Kanade, “Neural network-based\nPROCEEDINGS OF THE IEEE, VOL. 86, NO. 11, NOVEMBER 1998\n\n\f[97] [98] [99]\n[100] [101] [102]\n[103] [104] [105] [106] [107] [108] [109] [110] [111] [112] [113] [114] [115] [116] [117] [118] [119]\n\nface detection,” in Proc. IEEE CVPR’96, pp. 203–208. E. Osuna, R. Freund, and F. Girosi, “Training support vector machines: An application to face detection,” in Proc. IEEE CVPR’96, pp. 130–136. H. Bourlard and C. J. Wellekens, “Links between Markov models and multilayer perceptrons,” in Advances in Neural Information Processing Systems, D. Touretzky, Ed. Denver: Morgan-Kaufmann, vol. 1, 1989, pp. 186–187. Y. Bengio, R. De Mori, G. Flammia, and R. Kompe, “Neural network—Gaussian mixture hybrid for speech recognition or density estimation,” in Advances in Neural Information Processing Systems 4, J. E. Moody, S. J. Hanson, and R. P. Lippmann, Eds. Denver, CO: Morgan Kaufmann, 1992, pp. 175–182. F. C. N. Pereira and M. Riley, “Speech recognition by composition of weighted ﬁnite automata,” in Finite-State Devices for Natural Lague Processing. Cambridge, MA: MIT, 1997. M. Mohri, “Finite-state transducers in language and speech processing,” Computational Linguistics, vol. 23, no. 2, pp. 269–311, 1997. I. Guyon, M. Schenkel, and J. Denker, “Overview and synthesis of on-line cursive handwriting recognition techniques,” in Handbook on Optical Character Recognition and Document Image Analysis, P. S. P. Wang and H. Bunke, Eds. New York: World Scientiﬁc, 1996. M. Mohri and M. Riley, “Weighted determinization and minimization for large vocabulary recognition,” in Proc. Eurospeech ’97, Rhodes, Greece, pp. 131–134. Y. Bengio and P. Frasconi, “An input/output HMM architecture,” in Advances in Neural Information Processing Systems, vol. 7, G. Tesauro, D. Touretzky, and T. Leen, Eds. Cambridge, MA: MIT, pp. 427–434, 1996.\n, “Input/output HMM’s for sequence processing,” IEEE Trans. Neural Networks, vol. 7, no. 5, pp. 1231–1249, 1996. M. Mohri, F. C. N. Pereira, and M. Riley, A Rational Design for a Weighted Finite-State Transducer Library (Lecture Notes in Computer Science). New York: Springer Verlag, 1997. M. Rahim, C. H. Lee, and B. H. Juang, “Discriminative utterance veriﬁcation for connected digits recognition,” IEEE Trans. Speech Audio Processing, vol. 5, pp. 266–277, 1997. M. Rahim, Y. Bengio, and Y. LeCun, “Discriminative feature and model design for automatic speech recognition,” in Proc. Eurospeech ’97, Rhodes, Greece. S. Bengio and Y. Bengio, “An EM algorithm for asynchronous input/output hidden Markov models,” in Proc. International Conference on Neural Information Processing, Hong-King, 1996, pp. 328–334. C. Tappert, C. Suen, and T. Wakahara, “The state of the art in on-line handwriting recognition,” IEEE Trans. Pattern Anal. Machine Intell., vol. 8, pp. 787–808, Dec. 1990. S. Manke and U. Bodenhausen, “A connectionist recognizer for on-line cursive handwriting recognition,” in Proc. Int. Conf. Acoustics, Speech, Signal Processing, Adelaide, vol. 2, 1994, pp. 633–636. M. Gilloux and M. Leroux, “Recognition of cursive script amounts on postal checks,” in Proc. Europ. Conf. Postal Technol., Nantes, France, June 1993, pp. 705–712. D. Guillevic and C. Y. Suen, “Cursive script recognition applied to the processing of bank checks,” in Proc. Int. Conf. Document Analysis Recognition, Montreal, Canada, Aug. 1995, pp. 11–14. L. Lam, C. Y. Suen, D. Guillevic, N. W. Strathy, M. Cheriet, K. Liu, and J. N. Said, “Automatic processing of information on checks,” in Int. Conf. Systems, Man, and Cybernetics, Vancouver, Canada, Oct. 1995, pp. 2353–2358. C. J. C. Burges, J. I. Ben, J. S. Denker, Y. LeCun, and C. R. Nohl, “Off line recognition of handwritten postal words using neural networks,” Int. J. Pattern Recognit. Artiﬁcial Intell., vol. 7, no. 4, p. 689, 1993. Y. LeCun, Y. Bengio, D. Henderson, A. Weisbuch, H. Weissman, and L. Jackel, “On-line handwriting recognition with neural networks: Spatial representation versus temporal representation,” in Proc. Int. Conf. Handwriting Drawing, 1993. U. Mu¨ller, A. Gunzinger, and W. Guggenbu¨hl, “Fast neural net simulation with a DSP processor array,” IEEE Trans. Neural Networks, vol. 6, pp. 203–213, Jan. 1995. R. Battiti, “First- and second-order methods for learning: Between steepest descent and Newton’s method,” Neural Computation, vol. 4, no. 2, pp. 141–166, 1992. A. H. Kramer and A. Sangiovanni-Vincentelli, “Efﬁcient par-\n\n[120] [121]\n\nallel learning algorithms for neural networks,” in Advances in Neural Information Processing Systems, vol. 1, D. S. Touretzky, Ed. San Mateo, CA: Morgan Kaufmann, 1988, pp. 40–48. M. Moller, Efﬁcient Training of Feed-Forward Neural Networks, Ph.D. dissertation, Aarhus Univ., Aarhus, Denmark, 1993. S. Becker and Y. LeCun, “Improving the convergence of back-propagation learning with second-order methods,” Univ. Toronto Connectionist Res. Group, Toronto, Ontario, Canada, Tech. Rep. CRG-TR-88-5, Sept. 1988."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified.",
        "Readable original article obtained from a public mirror; the author-hosted PDF had a broken text encoding. Number-to-entry alignment is unreliable in parts of this extraction."
      ]
    },
    {
      "Slug": "deep-belief-nets",
      "Paper": "A Fast Learning Algorithm for Deep Belief Nets",
      "AtlasYear": 2006,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://www.cs.toronto.edu/~hinton/absps/fastnc.pdf",
      "PdfSha256": "8B7484127488743E79AE9DEC711694D0852EEC0B6B092922C4FAE2C43649B548",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 621,
          "EndLine": 643,
          "PdfPages": [
            10,
            11
          ],
          "Text": "Belongie, S., Malik, J., and Puzicha, J. (2002). Shape matching and object recognition using shape contexts. IEEE Transactions on Pattern Analysis and Machine Intelligence, 24(4):509–522.\nCarreira-Perpinan, M. A. and Hinton, G. E. (2005). On contrastive divergence learning. In Artiﬁcial Intelligence and Statistics, 2005.\nDecoste, D. and Schoelkopf, B. (2002). Training invariant support vector machines. Machine Learning, 46:161–190.\nFreund, Y. (1995). Boosting a weak learning algorithm by majority. Information and Computation, 12(2):256 – 285.\nFriedman, J. and Stuetzle, W. (1981). Projection pursuit regression. Journal of the American Statistical Association, 76:817–823.\nHinton, G. E. (2002). Training products of experts by minimizing contrastive divergence. Neural Computation, 14(8):1711–1800.\nHinton, G. E., Dayan, P., Frey, B. J., and Neal, R. (1995). The wake-sleep algorithm for self-organizing neural networks. Science, 268:1158–1161.\nLeCun, Y., Bottou, L., and Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324.\nLee, T. S. and Mumford, D. (2003). Hierarchical bayesian inference in the visual cortex. Journal of the Optical Society of America, A., 20:1434–1448.\nMarks, T. K. and Movellan, J. R. (2001). Diffusion networks, product of experts, and factor analysis. In Proc. Int. Conf. on Independent Component Analysis, pages 481–485.\nMayraz, G. and Hinton, G. E. (2001). Recognizing handwritten digits using hierarchical products of experts. IEEE Transactions on Pattern Analysis and Machine Intelligence, 24:189–197.\nNeal, R. (1992). Connectionist learning of belief networks. Artiﬁcial Intelligence, 56:71–113.\nNeal, R. M. and Hinton, G. E. (1998). A new view of the EM algorithm that justiﬁes incremental, sparse and other variants. In Jordan, M. I., editor, Learning in Graphical Models, pages 355—368. Kluwer Academic Publishers.\nNing, F., Delhomme, D., LeCun, Y., Piano, F., Bottou, L., and Barbano, P. (2005). Toward automatic phenotyping of developing embryos from videos. IEEE Transactions on Image Processing, 14(9):1360–1371.\nPearl, J. (1988). Probabilistic Inference in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann, San Mateo, CA.\n\n\fRoth, S. and Black, M. J. (2005). Fields of experts: A framework for learning image priors. In IEEE Conf. on Computer Vision and Pattern Recognition.\nSanger, T. D. (1989). Optimal unsupervised learning in a single-layer linear feedforward neural. Neural Networks, 2(6):459–473.\nSimard, P. Y., Steinkraus, D., and Platt, J. (2003). Best practice for convolutional neural networks applied to visual document analysis. In International Conference on Document Analysis and Recogntion (ICDAR), IEEE Computer Society, Los Alamitos, pages 958–962.\nTeh, Y. and Hinton, G. E. (2001). Rate-coded restricted Boltzmann machines for face recognition. In Advances in Neural Information Processing Systems, volume 13.\nTeh, Y., Welling, M., Osindero, S., and Hinton, G. E. (2003). Energy-based models for sparse overcomplete representations. Journal of Machine Learning Research, 4:1235– 1260.\nWelling, M., Hinton, G., and Osindero, S. (2003). Learning sparse topographic representations with products of Student-t distributions. In S. Becker, S. T. and Obermayer, K., editors, Advances in Neural Information Processing Systems 15, pages 1359–1366. MIT Press, Cambridge, MA.\nWelling, M., Rosen-Zvi, M., and Hinton, G. E. (2005). Exponential family harmoniums with an application to information retrieval. In Advances in Neural Information Processing Systems 17, pages 1481–1488. MIT Press, Cambridge, MA."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "imagenet",
      "Paper": "ImageNet: A Large-Scale Hierarchical Image Database",
      "AtlasYear": 2009,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://image-net.org/static_files/papers/imagenet_cvpr09.pdf",
      "PdfSha256": "9875AE02D26AD1C77554BCD6921610189D86C2609D542DE0FA43E34B0990F3B3",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 557,
          "EndLine": 562,
          "PdfPages": [
            8,
            9
          ],
          "Text": "[1] http://www.hunch.net/˜jl/. [2] The Chinese WordNet. http://bow.sinica.edu.tw. [3] The Spanish WordNet. http://www.lsi.upc.edu/˜nlp. [4] A. Artale, B. Magnini, and S. C. Wordnet for italian and its use for\nlexical discrimination. In AI*IA97, pages 16–19, 1997. [5] O. Boiman, E. Shechtman, and M. Irani. In defense of nearest-\nneighbor based image classiﬁcation. In CVPR08, pages 1–8, 2008. [6] B. Collins, J. Deng, K. Li, and L. Fei-Fei. Towards scalable dataset\nconstruction: An active learning approach. In ECCV08, pages I: 86– 98, 2008. [7] M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2008 (VOC2008) Results. http://www.pascal-network.org/ challenges/VOC/voc2008/workshop/. [8] L. Fei-Fei, R. Fergus, and P. Perona. One-shot learning of object categories. PAMI, 28(4):594–611, April 2006. [9] C. Fellbaum. WordNet: An Electronic Lexical Database. Bradford Books, 1998. [10] R. Fergus, L. Fei-Fei, P. Perona, and A. Zisserman. Learning object categories from google’s image search. In ICCV05, pages II: 1816– 1823, 2005. [11] M. Fink and S. Ullman. From aardvark to zorro: A benchmark for mammal image classiﬁcation. IJCV, 77(1-3):143–156, May 2008. [12] G. Grifﬁn, A. Holub, and P. Perona. Caltech-256 object category dataset. Technical Report 7694, Caltech, 2007. [13] G. Huang, M. Ramesh, T. Berg, and E. Learned Miller. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. Technical Report 07-49, UMass, 2007. [14] L.-J. Li, G. Wang, and L. Fei-Fei. OPTIMOL: automatic Online Picture collecTion via Incremental MOdel Learning. In CVPR07, pages 1–8, 2007. [15] D. Lowe. Distinctive image features from scale-invariant keypoints. IJCV, 60(2):91–110, November 2004. [16] M. Marszalek and C. Schmid. Semantic hierarchies for visual object recognition. In CVPR07, pages 1–7, 2007. [17] M. Marszalek and C. Schmid. Constructing category hierarchies for visual recognition. In ECCV08, pages IV: 479–491, 2008. [18] D. Nister and H. Stewenius. Scalable recognition with a vocabulary tree. In CVPR06, pages II: 2161–2168, 2006. [19] P. Phillips, H. Wechsler, J. Huang, and P. Rauss. The feret database and evaluation procedure for face-recognition algorithms. IVC, 16(5):295–306, April 1998. [20] E. Rosch and B. Lloyd. Principles of categorization. In Cognition and categorization, pages 27–48, 1978. [21] B. Russell, A. Torralba, K. Murphy, and W. Freeman. Labelme: A database and web-based tool for image annotation. IJCV, 77(13):157–173, May 2008. [22] J. Shotton, J. Winn, C. Rother, and A. Criminisi. Textonboost: Joint appearance, shape and context modeling for multi-class object recognition and segmentation. In ECCV06, pages I: 1–15, 2006. [23] A. Sorokin and D. Forsyth. Utility data annotation with amazon mechanical turk. In InterNet08, pages 1–8, 2008. [24] A. Torralba, R. Fergus, and W. Freeman. 80 million tiny images: A large data set for nonparametric object and scene recognition. PAMI, 30(11):1958–1970, November 2008. [25] L. von Ahn and L. Dabbish. Labeling images with a computer game. In CHI04, pages 319–326, 2004. [26] P. Vossen, K. Hofmann, M. de Rijke, E. Tjong Kim Sang, and K. Deschacht. The Cornetto database: Architecture and user-scenarios. In Proceedings DIR 2007, pages 89–96, 2007. [27] B. Yao, X. Yang, and S. Zhu. Introduction to a large-scale general purpose ground truth database: Methodology, annotation tool and benchmarks. In EMMCVPR07, pages 169–183, 2007. [28] A. Zweig and D. Weinshall. Exploiting object hierarchy: Combining models from different category levels. In ICCV07, pages 1–8, 2007."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "scikit-learn",
      "Paper": "Scikit-learn: Machine Learning in Python",
      "AtlasYear": 2011,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://jmlr.org/papers/volume12/pedregosa11a/pedregosa11a.pdf",
      "PdfSha256": "1C338A6B3C6C1DCAFDA3990A8098B8B50DD6C77616361396FB75F6822C6E0778",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 115,
          "EndLine": 134,
          "PdfPages": [
            5,
            6,
            7
          ],
          "Text": "D. Albanese, G. Merler, S.and Jurman, and R. Visintainer. MLPy: high-performance python package for predictive modeling. In NIPS, MLOSS Workshop, 2008.\nC.C. Chang and C.J. Lin. LIBSVM: a library for support vector machines. http://www.csie. ntu.edu.tw/cjlin/libsvm, 2001.\nP.F. Dubois, editor. Python: Batteries Included, volume 9 of Computing in Science & Engineering. IEEE/AIP, May 2007.\nR.E. Fan, K.W. Chang, C.J. Hsieh, X.R. Wang, and C.J. Lin. LIBLINEAR: a library for large linear classiﬁcation. The Journal of Machine Learning Research, 9:1871–1874, 2008.\nJ. Friedman, T. Hastie, and R. Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1):1, 2010.\nI Guyon, S. R. Gunn, A. Ben-Hur, and G. Dror. Result analysis of the NIPS 2003 feature selection challenge, 2004.\nM. Hanke, Y.O. Halchenko, P.B. Sederberg, S.J. Hanson, J.V. Haxby, and S. Pollmann. PyMVPA: A Python toolbox for multivariate pattern analysis of fMRI data. Neuroinformatics, 7(1):37–53, 2009.\n2829\n\n\fPEDREGOSA, VAROQUAUX, GRAMFORT ET AL.\nT. Hastie and B. Efron. Least Angle Regression, Lasso and Forward Stagewise. http://cran. r-project.org/web/packages/lars/lars.pdf, 2004.\nV. Michel, A. Gramfort, G. Varoquaux, E. Eger, C. Keribin, and B. Thirion. A supervised clustering approach for fMRI-based inference of brain states. Patt Rec, page epub ahead of print, April 2011. doi: 10.1016/j.patcog.2011.04.006.\nK.J. Milmann and M. Avaizis, editors. Scientiﬁc Python, volume 11 of Computing in Science & Engineering. IEEE/AIP, March 2011.\nS.M. Omohundro. Five balltree construction algorithms. ICSI Technical Report TR-89-063, 1989. V. Rokhlin, A. Szlam, and M. Tygert. A randomized algorithm for principal component analysis.\nSIAM Journal on Matrix Analysis and Applications, 31(3):1100–1124, 2009. T. Schaul, J. Bayer, D. Wierstra, Y. Sun, M. Felder, F. Sehnke, T. Ru¨ckstieß, and J. Schmidhuber.\nPyBrain. The Journal of Machine Learning Research, 11:743–746, 2010. S. Sonnenburg, G. Ra¨tsch, S. Henschel, C. Widmer, J. Behr, A. Zien, F. de Bona, A. Binder, C. Gehl,\nand V. Franc. The SHOGUN machine learning toolbox. Journal of Machine Learning Research, 11:1799–1802, 2010. S. Van der Walt, S.C Colbert, and G. Varoquaux. The NumPy array: A structure for efﬁcient numerical computation. Computing in Science and Engineering, 11, 2011. T. Zito, N. Wilbert, L. Wiskott, and P. Berkes. Modular toolkit for data processing (MDP): A Python data processing framework. Frontiers in Neuroinformatics, 2, 2008.\n2830"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "alexnet",
      "Paper": "ImageNet Classification with Deep Convolutional Neural Networks",
      "AtlasYear": 2012,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://papers.nips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf",
      "PdfSha256": "90137160C57217953D5F61857E64CA58E85F06E1B13B4F475C918B1B582B9771",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 291,
          "EndLine": 301,
          "PdfPages": [
            9,
            10
          ],
          "Text": "[1] R.M. Bell and Y. Koren. Lessons from the netﬂix prize challenge. ACM SIGKDD Explorations Newsletter, 9(2):75–79, 2007.\n[2] A. Berg, J. Deng, and L. Fei-Fei. Large scale visual recognition challenge 2010. www.imagenet.org/challenges. 2010.\n[3] L. Breiman. Random forests. Machine learning, 45(1):5–32, 2001. [4] D. Cires¸an, U. Meier, and J. Schmidhuber. Multi-column deep neural networks for image classiﬁcation.\nArxiv preprint arXiv:1202.2745, 2012. [5] D.C. Cires¸an, U. Meier, J. Masci, L.M. Gambardella, and J. Schmidhuber. High-performance neural\nnetworks for visual object classiﬁcation. Arxiv preprint arXiv:1102.0183, 2011. [6] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical\nImage Database. In CVPR09, 2009. [7] J. Deng, A. Berg, S. Satheesh, H. Su, A. Khosla, and L. Fei-Fei. ILSVRC-2012, 2012. URL\nhttp://www.image-net.org/challenges/LSVRC/2012/. [8] L. Fei-Fei, R. Fergus, and P. Perona. Learning generative visual models from few training examples: An\nincremental bayesian approach tested on 101 object categories. Computer Vision and Image Understanding, 106(1):59–70, 2007. [9] G. Grifﬁn, A. Holub, and P. Perona. Caltech-256 object category dataset. Technical Report 7694, California Institute of Technology, 2007. URL http://authors.library.caltech.edu/7694. [10] G.E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R.R. Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv:1207.0580, 2012. [11] K. Jarrett, K. Kavukcuoglu, M. A. Ranzato, and Y. LeCun. What is the best multi-stage architecture for object recognition? In International Conference on Computer Vision, pages 2146–2153. IEEE, 2009. [12] A. Krizhevsky. Learning multiple layers of features from tiny images. Master’s thesis, Department of Computer Science, University of Toronto, 2009. [13] A. Krizhevsky. Convolutional deep belief networks on cifar-10. Unpublished manuscript, 2010. [14] A. Krizhevsky and G.E. Hinton. Using very deep autoencoders for content-based image retrieval. In ESANN, 2011. [15] Y. Le Cun, B. Boser, J.S. Denker, D. Henderson, R.E. Howard, W. Hubbard, L.D. Jackel, et al. Handwritten digit recognition with a back-propagation network. In Advances in neural information processing systems, 1990. [16] Y. LeCun, F.J. Huang, and L. Bottou. Learning methods for generic object recognition with invariance to pose and lighting. In Computer Vision and Pattern Recognition, 2004. CVPR 2004. Proceedings of the 2004 IEEE Computer Society Conference on, volume 2, pages II–97. IEEE, 2004. [17] Y. LeCun, K. Kavukcuoglu, and C. Farabet. Convolutional networks and applications in vision. In Circuits and Systems (ISCAS), Proceedings of 2010 IEEE International Symposium on, pages 253–256. IEEE, 2010. [18] H. Lee, R. Grosse, R. Ranganath, and A.Y. Ng. Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations. In Proceedings of the 26th Annual International Conference on Machine Learning, pages 609–616. ACM, 2009. [19] T. Mensink, J. Verbeek, F. Perronnin, and G. Csurka. Metric Learning for Large Scale Image Classiﬁcation: Generalizing to New Classes at Near-Zero Cost. In ECCV - European Conference on Computer Vision, Florence, Italy, October 2012. [20] V. Nair and G. E. Hinton. Rectiﬁed linear units improve restricted boltzmann machines. In Proc. 27th International Conference on Machine Learning, 2010. [21] N. Pinto, D.D. Cox, and J.J. DiCarlo. Why is real-world visual object recognition hard? PLoS computational biology, 4(1):e27, 2008. [22] N. Pinto, D. Doukhan, J.J. DiCarlo, and D.D. Cox. A high-throughput screening approach to discovering good forms of biologically inspired visual representation. PLoS computational biology, 5(11):e1000579, 2009. [23] B.C. Russell, A. Torralba, K.P. Murphy, and W.T. Freeman. Labelme: a database and web-based tool for image annotation. International journal of computer vision, 77(1):157–173, 2008. [24] J. Sánchez and F. Perronnin. High-dimensional signature compression for large-scale image classiﬁcation. In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pages 1665–1672. IEEE, 2011. [25] P.Y. Simard, D. Steinkraus, and J.C. Platt. Best practices for convolutional neural networks applied to visual document analysis. In Proceedings of the Seventh International Conference on Document Analysis and Recognition, volume 2, pages 958–962, 2003. [26] S.C. Turaga, J.F. Murray, V. Jain, F. Roth, M. Helmstaedter, K. Briggman, W. Denk, and H.S. Seung. Convolutional networks can learn to generate afﬁnity graphs for image segmentation. Neural Computation, 22(2):511–538, 2010.\n9"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "word2vec",
      "Paper": "Efficient Estimation of Word Representations in Vector Space",
      "AtlasYear": 2013,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/1301.3781",
      "PdfSha256": "A44D7E22D2005752271C9CC1929C6462D4C8270916B063977992A883E3A54362",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 523,
          "EndLine": 560,
          "PdfPages": [
            11,
            12,
            13
          ],
          "Text": "[1] Y. Bengio, R. Ducharme, P. Vincent. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137-1155, 2003.\n[2] Y. Bengio, Y. LeCun. Scaling learning algorithms towards AI. In: Large-Scale Kernel Machines, MIT Press, 2007.\n[3] T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean. Large language models in machine translation. In Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Computational Language Learning, 2007.\n[4] R. Collobert and J. Weston. A Uniﬁed Architecture for Natural Language Processing: Deep Neural Networks with Multitask Learning. In International Conference on Machine Learning, ICML, 2008.\n[5] R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu and P. Kuksa. Natural Language Processing (Almost) from Scratch. Journal of Machine Learning Research, 12:24932537, 2011.\n[6] J. Dean, G.S. Corrado, R. Monga, K. Chen, M. Devin, Q.V. Le, M.Z. Mao, M.A. Ranzato, A. Senior, P. Tucker, K. Yang, A. Y. Ng., Large Scale Distributed Deep Networks, NIPS, 2012.\n[7] J.C. Duchi, E. Hazan, and Y. Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 2011.\n[8] J. Elman. Finding Structure in Time. Cognitive Science, 14, 179-211, 1990.\n[9] Eric H. Huang, R. Socher, C. D. Manning and Andrew Y. Ng. Improving Word Representations via Global Context and Multiple Word Prototypes. In: Proc. Association for Computational Linguistics, 2012.\n[10] G.E. Hinton, J.L. McClelland, D.E. Rumelhart. Distributed representations. In: Parallel distributed processing: Explorations in the microstructure of cognition. Volume 1: Foundations, MIT Press, 1986.\n[11] D.A. Jurgens, S.M. Mohammad, P.D. Turney, K.J. Holyoak. Semeval-2012 task 2: Measuring degrees of relational similarity. In: Proceedings of the 6th International Workshop on Semantic Evaluation (SemEval 2012), 2012.\n[12] A.L. Maas, R.E. Daly, P.T. Pham, D. Huang, A.Y. Ng, and C. Potts. Learning word vectors for sentiment analysis. In Proceedings of ACL, 2011.\n[13] T. Mikolov. Language Modeling for Speech Recognition in Czech, Masters thesis, Brno University of Technology, 2007.\n[14] T. Mikolov, J. Kopecky´, L. Burget, O. Glembek and J. Cˇ ernocky´. Neural network based language models for higly inﬂective languages, In: Proc. ICASSP 2009.\n[15] T. Mikolov, M. Karaﬁa´t, L. Burget, J. Cˇ ernocky´, S. Khudanpur. Recurrent neural network based language model, In: Proceedings of Interspeech, 2010.\n[16] T. Mikolov, S. Kombrink, L. Burget, J. Cˇ ernocky´, S. Khudanpur. Extensions of recurrent neural network language model, In: Proceedings of ICASSP 2011.\n[17] T. Mikolov, A. Deoras, S. Kombrink, L. Burget, J. Cˇ ernocky´. Empirical Evaluation and Combination of Advanced Language Modeling Techniques, In: Proceedings of Interspeech, 2011.\n4The code is available at https://code.google.com/p/word2vec/\n11\n\n\f[18] T. Mikolov, A. Deoras, D. Povey, L. Burget, J. Cˇ ernocky´. Strategies for Training Large Scale Neural Network Language Models, In: Proc. Automatic Speech Recognition and Understanding, 2011.\n[19] T. Mikolov. Statistical Language Models based on Neural Networks. PhD thesis, Brno University of Technology, 2012.\n[20] T. Mikolov, W.T. Yih, G. Zweig. Linguistic Regularities in Continuous Space Word Representations. NAACL HLT 2013.\n[21] T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean. Distributed Representations of Words and Phrases and their Compositionality. Accepted to NIPS 2013.\n[22] A. Mnih, G. Hinton. Three new graphical models for statistical language modelling. ICML, 2007.\n[23] A. Mnih, G. Hinton. A Scalable Hierarchical Distributed Language Model. Advances in Neural Information Processing Systems 21, MIT Press, 2009.\n[24] A. Mnih, Y.W. Teh. A fast and simple algorithm for training neural probabilistic language models. ICML, 2012.\n[25] F. Morin, Y. Bengio. Hierarchical Probabilistic Neural Network Language Model. AISTATS, 2005.\n[26] D. E. Rumelhart, G. E. Hinton, R. J. Williams. Learning internal representations by backpropagating errors. Nature, 323:533.536, 1986.\n[27] H. Schwenk. Continuous space language models. Computer Speech and Language, vol. 21, 2007.\n[28] R. Socher, E.H. Huang, J. Pennington, A.Y. Ng, and C.D. Manning. Dynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase Detection. In NIPS, 2011.\n[29] J. Turian, L. Ratinov, Y. Bengio. Word Representations: A Simple and General Method for Semi-Supervised Learning. In: Proc. Association for Computational Linguistics, 2010.\n[30] P. D. Turney. Measuring Semantic Similarity by Latent Relational Analysis. In: Proc. International Joint Conference on Artiﬁcial Intelligence, 2005.\n[31] A. Zhila, W.T. Yih, C. Meek, G. Zweig, T. Mikolov. Combining Heterogeneous Models for Measuring Relational Similarity. NAACL HLT 2013.\n[32] G. Zweig, C.J.C. Burges. The Microsoft Research Sentence Completion Challenge, Microsoft Research Technical Report MSR-TR-2011-129, 2011.\n12"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "gans",
      "Paper": "Generative Adversarial Networks",
      "AtlasYear": 2014,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/1406.2661",
      "PdfSha256": "FF5819E3A7B713C3BD3107B7DE3D51FE0A347AA5D8444F0EFDCF2345EF0A8B63",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 380,
          "EndLine": 415,
          "PdfPages": [
            8,
            9,
            10
          ],
          "Text": "[1] Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012). Theano: new features and speed improvements. Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop.\n[2] Bengio, Y. (2009). Learning deep architectures for AI. Now Publishers.\n[3] Bengio, Y., Mesnil, G., Dauphin, Y., and Rifai, S. (2013a). Better mixing via deep representations. In ICML’13.\n[4] Bengio, Y., Yao, L., Alain, G., and Vincent, P. (2013b). Generalized denoising auto-encoders as generative models. In NIPS26. Nips Foundation.\n[5] Bengio, Y., Thibodeau-Laufer, E., and Yosinski, J. (2014a). Deep generative stochastic networks trainable by backprop. In ICML’14.\n[6] Bengio, Y., Thibodeau-Laufer, E., Alain, G., and Yosinski, J. (2014b). Deep generative stochastic networks trainable by backprop. In Proceedings of the 30th International Conference on Machine Learning (ICML’14).\n[7] Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010). Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientiﬁc Computing Conference (SciPy). Oral Presentation.\n[8] Breuleux, O., Bengio, Y., and Vincent, P. (2011). Quickly generating representative samples from an RBM-derived process. Neural Computation, 23(8), 2053–2073.\n[9] Glorot, X., Bordes, A., and Bengio, Y. (2011). Deep sparse rectiﬁer neural networks. In AISTATS’2011.\n[10] Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013a). Maxout networks. In ICML’2013.\n[11] Goodfellow, I. J., Mirza, M., Courville, A., and Bengio, Y. (2013b). Multi-prediction deep Boltzmann machines. In NIPS’2013.\n[12] Goodfellow, I. J., Warde-Farley, D., Lamblin, P., Dumoulin, V., Mirza, M., Pascanu, R., Bergstra, J., Bastien, F., and Bengio, Y. (2013c). Pylearn2: a machine learning research library. arXiv preprint arXiv:1308.4214.\n[13] Gutmann, M. and Hyvarinen, A. (2010). Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In AISTATS’2010.\n[14] Hinton, G., Deng, L., Dahl, G. E., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., and Kingsbury, B. (2012a). Deep neural networks for acoustic modeling in speech recognition. IEEE Signal Processing Magazine, 29(6), 82–97.\n[15] Hinton, G. E., Dayan, P., Frey, B. J., and Neal, R. M. (1995). The wake-sleep algorithm for unsupervised neural networks. Science, 268, 1558–1161.\n8\n\n\f[16] Hinton, G. E., Osindero, S., and Teh, Y. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18, 1527–1554.\n[17] Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2012b). Improving neural networks by preventing co-adaptation of feature detectors. Technical report, arXiv:1207.0580.\n[18] Hyva¨rinen, A. (2005). Estimation of non-normalized statistical models using score matching. J. Machine Learning Res., 6.\n[19] Jarrett, K., Kavukcuoglu, K., Ranzato, M., and LeCun, Y. (2009). What is the best multi-stage architecture for object recognition? In Proc. International Conference on Computer Vision (ICCV’09), pages 2146–2153. IEEE.\n[20] Kingma, D. P. and Welling, M. (2014). Auto-encoding variational bayes. In Proceedings of the International Conference on Learning Representations (ICLR).\n[21] Krizhevsky, A. and Hinton, G. (2009). Learning multiple layers of features from tiny images. Technical report, University of Toronto.\n[22] Krizhevsky, A., Sutskever, I., and Hinton, G. (2012). ImageNet classiﬁcation with deep convolutional neural networks. In NIPS’2012.\n[23] LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278–2324.\n[24] Rezende, D. J., Mohamed, S., and Wierstra, D. (2014). Stochastic backpropagation and approximate inference in deep generative models. Technical report, arXiv:1401.4082.\n[25] Rifai, S., Bengio, Y., Dauphin, Y., and Vincent, P. (2012). A generative process for sampling contractive auto-encoders. In ICML’12.\n[26] Salakhutdinov, R. and Hinton, G. E. (2009). Deep Boltzmann machines. In AISTATS’2009, pages 448– 455.\n[27] Smolensky, P. (1986). Information processing in dynamical systems: Foundations of harmony theory. In D. E. Rumelhart and J. L. McClelland, editors, Parallel Distributed Processing, volume 1, chapter 6, pages 194–281. MIT Press, Cambridge.\n[28] Susskind, J., Anderson, A., and Hinton, G. E. (2010). The Toronto face dataset. Technical Report UTML TR 2010-001, U. Toronto.\n[29] Tieleman, T. (2008). Training restricted Boltzmann machines using approximations to the likelihood gradient. In W. W. Cohen, A. McCallum, and S. T. Roweis, editors, ICML 2008, pages 1064–1071. ACM.\n[30] Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. (2008). Extracting and composing robust features with denoising autoencoders. In ICML 2008.\n[31] Younes, L. (1999). On the convergence of Markovian stochastic algorithms with rapidly decreasing ergodicity rates. Stochastics and Stochastic Reports, 65(3), 177–228.\n9"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "dropout",
      "Paper": "Dropout: A Simple Way to Prevent Neural Networks from Overfitting",
      "AtlasYear": 2014,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://jmlr.org/papers/volume15/srivastava14a/srivastava14a.pdf",
      "PdfSha256": "9C196CCBE6C6A595A1ADBA6CD030D35F7C2E548BBF5E7F1278B0109D8DD9EBAA",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 952,
          "EndLine": 996,
          "PdfPages": [
            28,
            29,
            30,
            31
          ],
          "Text": "M. Chen, Z. Xu, K. Weinberger, and F. Sha. Marginalized denoising autoencoders for domain adaptation. In Proceedings of the 29th International Conference on Machine Learning, pages 767–774. ACM, 2012.\nG. E. Dahl, M. Ranzato, A. Mohamed, and G. E. Hinton. Phone recognition with the meancovariance restricted Boltzmann machine. In Advances in Neural Information Processing Systems 23, pages 469–477, 2010.\nO. Dekel, O. Shamir, and L. Xiao. Learning to classify with missing and corrupted features. Machine Learning, 81(2):149–178, 2010.\nA. Globerson and S. Roweis. Nightmare at test time: robust learning by feature deletion. In Proceedings of the 23rd International Conference on Machine Learning, pages 353–360. ACM, 2006.\nI. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio. Maxout networks. In Proceedings of the 30th International Conference on Machine Learning, pages 1319– 1327. ACM, 2013.\nG. Hinton and R. Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313(5786):504 – 507, 2006.\nG. E. Hinton, S. Osindero, and Y. Teh. A fast learning algorithm for deep belief nets. Neural Computation, 18:1527–1554, 2006.\nK. Jarrett, K. Kavukcuoglu, M. Ranzato, and Y. LeCun. What is the best multi-stage architecture for object recognition? In Proceedings of the International Conference on Computer Vision (ICCV’09). IEEE, 2009.\nA. Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.\nA. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classiﬁcation with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25, pages 1106–1114, 2012.\nY. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1(4):541–551, 1989.\nY. Lin, F. Lv, S. Zhu, M. Yang, T. Cour, K. Yu, L. Cao, Z. Li, M.-H. Tsai, X. Zhou, T. Huang, and T. Zhang. Imagenet classiﬁcation: fast descriptor coding and large-scale svm training. Large scale visual recognition challenge, 2010.\nA. Livnat, C. Papadimitriou, N. Pippenger, and M. W. Feldman. Sex, mixability, and modularity. Proceedings of the National Academy of Sciences, 107(4):1452–1457, 2010.\nV. Mnih. CUDAMat: a CUDA-based matrix class for Python. Technical Report UTML TR 2009-004, Department of Computer Science, University of Toronto, November 2009.\n1956\n\n\fDropout\nA. Mohamed, G. E. Dahl, and G. E. Hinton. Acoustic modeling using deep belief networks. IEEE Transactions on Audio, Speech, and Language Processing, 2010.\nR. M. Neal. Bayesian Learning for Neural Networks. Springer-Verlag New York, Inc., 1996.\nY. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011.\nS. J. Nowlan and G. E. Hinton. Simplifying neural networks by soft weight-sharing. Neural Computation, 4(4), 1992.\nD. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely. The Kaldi Speech Recognition Toolkit. In IEEE 2011 Workshop on Automatic Speech Recognition and Understanding. IEEE Signal Processing Society, 2011.\nR. Salakhutdinov and G. Hinton. Deep Boltzmann machines. In Proceedings of the International Conference on Artiﬁcial Intelligence and Statistics, volume 5, pages 448–455, 2009.\nR. Salakhutdinov and A. Mnih. Bayesian probabilistic matrix factorization using Markov chain Monte Carlo. In Proceedings of the 25th International Conference on Machine Learning. ACM, 2008.\nJ. Sanchez and F. Perronnin. High-dimensional signature compression for large-scale image classiﬁcation. In Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition, pages 1665–1672, 2011.\nP. Sermanet, S. Chintala, and Y. LeCun. Convolutional neural networks applied to house numbers digit classiﬁcation. In International Conference on Pattern Recognition (ICPR 2012), 2012.\nP. Simard, D. Steinkraus, and J. Platt. Best practices for convolutional neural networks applied to visual document analysis. In Proceedings of the Seventh International Conference on Document Analysis and Recognition, volume 2, pages 958–962, 2003.\nJ. Snoek, H. Larochelle, and R. Adams. Practical Bayesian optimization of machine learning algorithms. In Advances in Neural Information Processing Systems 25, pages 2960–2968, 2012.\nN. Srebro and A. Shraibman. Rank, trace-norm and max-norm. In Proceedings of the 18th annual conference on Learning Theory, COLT’05, pages 545–560. Springer-Verlag, 2005.\nN. Srivastava. Improving Neural Networks with Dropout. Master’s thesis, University of Toronto, January 2013.\nR. Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society. Series B. Methodological, 58(1):267–288, 1996.\n1957\n\n\fSrivastava, Hinton, Krizhevsky, Sutskever and Salakhutdinov\nA. N. Tikhonov. On the stability of inverse problems. Doklady Akademii Nauk SSSR, 39(5): 195–198, 1943.\nL. van der Maaten, M. Chen, S. Tyree, and K. Q. Weinberger. Learning with marginalized corrupted features. In Proceedings of the 30th International Conference on Machine Learning, pages 410–418. ACM, 2013.\nP. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th International Conference on Machine Learning, pages 1096–1103. ACM, 2008.\nP. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. In Proceedings of the 27th International Conference on Machine Learning, pages 3371–3408. ACM, 2010.\nS. Wager, S. Wang, and P. Liang. Dropout training as adaptive regularization. In Advances in Neural Information Processing Systems 26, pages 351–359, 2013.\nS. Wang and C. D. Manning. Fast dropout training. In Proceedings of the 30th International Conference on Machine Learning, pages 118–126. ACM, 2013.\nH. Y. Xiong, Y. Barash, and B. J. Frey. Bayesian prediction of tissue-regulated splicing using RNA sequence and cellular context. Bioinformatics, 27(18):2554–2562, 2011.\nM. D. Zeiler and R. Fergus. Stochastic pooling for regularization of deep convolutional neural networks. CoRR, abs/1301.3557, 2013.\n1958"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "deep-q-learning",
      "Paper": "Human-level control through deep reinforcement learning",
      "AtlasYear": 2015,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://web.stanford.edu/class/psych209/Readings/MnihEtAlHassibis15NatureControlDeepRL.pdf",
      "PdfSha256": "7E76CFD09E121CF546992FFC5B2116C0DA73EC1595069F142A29DA5BAE400F4E",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 340,
          "EndLine": 373,
          "PdfPages": [
            4,
            5
          ],
          "Text": "1. Sutton, R. & Barto, A. Reinforcement Learning: An Introduction (MIT Press, 1998). 2. Thorndike, E. L. Animal Intelligence: Experimental studies (Macmillan, 1911). 3. Schultz, W., Dayan, P. & Montague, P. R. A neural substrate of prediction and\nreward. Science 275, 1593–1599 (1997). 4. Serre, T., Wolf, L. & Poggio, T. Object recognition with features inspired by visual\ncortex. Proc. IEEE. Comput. Soc. Conf. Comput. Vis. Pattern. Recognit. 994–1000 (2005). 5. Fukushima, K. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biol. Cybern. 36, 193–202 (1980).\n\n532 | NATURE | VOL 518 | 26 FEBRUARY 2015\n©2015 Macmillan Publishers Limited. All rights reserved\n\n\fLETTER RESEARCH\n\n6. Tesauro, G. Temporal difference learning and TD-Gammon. Commun. ACM 38, 58–68 (1995).\n7. Riedmiller, M., Gabel, T., Hafner, R. & Lange, S. Reinforcement learning for robot soccer. Auton. Robots 27, 55–73 (2009).\n8. Diuk, C., Cohen, A. & Littman, M. L. An object-oriented representation for efficient reinforcement learning. Proc. Int. Conf. Mach. Learn. 240–247 (2008).\n9. Bengio, Y. Learning deep architectures for AI. Foundations and Trends in Machine Learning 2, 1–127 (2009).\n10. Krizhevsky, A., Sutskever, I. & Hinton, G. ImageNet classification with deep convolutional neural networks. Adv.Neural Inf.Process.Syst.25, 1106–1114 (2012).\n11. Hinton, G. E. & Salakhutdinov, R. R. Reducing the dimensionality of data with neural networks. Science 313, 504–507 (2006).\n12. Bellemare, M. G., Naddaf, Y., Veness, J. & Bowling, M. The arcade learning environment: An evaluation platform for general agents. J. Artif. Intell. Res. 47, 253–279 (2013).\n13. Legg, S. & Hutter, M. Universal Intelligence: a definition of machine intelligence. Minds Mach. 17, 391–444 (2007).\n14. Genesereth, M., Love, N. & Pell, B. General game playing: overview of the AAAI competition. AI Mag. 26, 62–72 (2005).\n15. Bellemare, M. G., Veness, J. & Bowling, M. Investigating contingency awareness using Atari 2600 games. Proc. Conf. AAAI. Artif. Intell. 864–871 (2012).\n16. McClelland, J. L., Rumelhart, D. E. & Group, T. P. R. Parallel Distributed Processing: Explorations in the Microstructure of Cognition (MIT Press, 1986).\n17. LeCun, Y., Bottou, L., Bengio, Y. & Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 86, 2278–2324 (1998).\n18. Hubel, D. H. & Wiesel, T. N. Shape and arrangement of columns in cat’s striate cortex. J. Physiol. 165, 559–568 (1963).\n19. Watkins, C. J. & Dayan, P. Q-learning. Mach. Learn. 8, 279–292 (1992). 20. Tsitsiklis, J. & Roy, B. V. An analysis of temporal-difference learning with function\napproximation. IEEE Trans. Automat. Contr. 42, 674–690 (1997). 21. McClelland, J. L., McNaughton, B. L. & O’Reilly, R. C. Why there are complementary\nlearning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory. Psychol. Rev. 102, 419–457 (1995). 22. O’Neill, J., Pleydell-Bouverie, B., Dupret, D. & Csicsvari, J. Play it again: reactivation of waking experience and memory. Trends Neurosci. 33, 220–229 (2010).\n\n23. Lin, L.-J. Reinforcement learning for robots using neural networks. Technical Report, DTIC Document (1993).\n24. Riedmiller, M. Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method. Mach. Learn.: ECML, 3720, 317–328 (Springer, 2005).\n25. Van der Maaten, L. J. P. & Hinton, G. E. Visualizing high-dimensional data using t-SNE. J. Mach. Learn. Res. 9, 2579–2605 (2008).\n26. Lange, S. & Riedmiller, M. Deep auto-encoder neural networks in reinforcement learning. Proc. Int. Jt. Conf. Neural. Netw. 1–8 (2010).\n27. Law, C.-T. & Gold, J. I. Reinforcement learning can account for associative and perceptual learning on a visual decision task. Nature Neurosci. 12, 655 (2009).\n28. Sigala, N. & Logothetis, N. K. Visual categorization shapes feature selectivity in the primate temporal cortex. Nature 415, 318–320 (2002).\n29. Bendor, D. & Wilson, M. A. Biasing the content of hippocampal replay during sleep. Nature Neurosci. 15, 1439–1444 (2012).\n30. Moore, A. & Atkeson, C. Prioritized sweeping: reinforcement learning with less data and less real time. Mach. Learn. 13, 103–130 (1993)."
        },
        {
          "Section": "Reference section 2",
          "StartLine": 817,
          "EndLine": 819,
          "PdfPages": [
            7
          ],
          "Text": "31. Jarrett, K., Kavukcuoglu, K., Ranzato, M. A. & LeCun, Y. What is the best multi-stage architecture for object recognition? Proc. IEEE. Int. Conf. Comput. Vis. 2146–2153 (2009).\n32. Nair, V. & Hinton, G. E. Rectified linear units improve restricted Boltzmann machines. Proc. Int. Conf. Mach. Learn. 807–814 (2010).\n33. Kaelbling, L. P., Littman, M. L. & Cassandra, A. R. Planning and acting in partially observable stochastic domains. Artificial Intelligence 101, 99–134 (1994)."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "resnet",
      "Paper": "Deep Residual Learning for Image Recognition",
      "AtlasYear": 2015,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/1512.03385",
      "PdfSha256": "1E0651B6810ECBA34A3DBC5B5B0209226F889004607C1F203540A48D64E5A93A",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 717,
          "EndLine": 748,
          "PdfPages": [
            9
          ],
          "Text": "[1] Y. Bengio, P. Simard, and P. Frasconi. Learning long-term dependencies with gradient descent is difﬁcult. IEEE Transactions on Neural Networks, 5(2):157–166, 1994.\n[2] C. M. Bishop. Neural networks for pattern recognition. Oxford university press, 1995.\n[3] W. L. Briggs, S. F. McCormick, et al. A Multigrid Tutorial. Siam, 2000.\n[4] K. Chatﬁeld, V. Lempitsky, A. Vedaldi, and A. Zisserman. The devil is in the details: an evaluation of recent feature encoding methods. In BMVC, 2011.\n[5] M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The Pascal Visual Object Classes (VOC) Challenge. IJCV, pages 303–338, 2010.\n[6] S. Gidaris and N. Komodakis. Object detection via a multi-region & semantic segmentation-aware cnn model. In ICCV, 2015.\n[7] R. Girshick. Fast R-CNN. In ICCV, 2015. [8] R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hier-\narchies for accurate object detection and semantic segmentation. In CVPR, 2014. [9] X. Glorot and Y. Bengio. Understanding the difﬁculty of training deep feedforward neural networks. In AISTATS, 2010. [10] I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio. Maxout networks. arXiv:1302.4389, 2013. [11] K. He and J. Sun. Convolutional neural networks at constrained time cost. In CVPR, 2015. [12] K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. In ECCV, 2014. [13] K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectiﬁers: Surpassing human-level performance on imagenet classiﬁcation. In ICCV, 2015. [14] G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov. Improving neural networks by preventing coadaptation of feature detectors. arXiv:1207.0580, 2012. [15] S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997. [16] S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. [17] H. Jegou, M. Douze, and C. Schmid. Product quantization for nearest neighbor search. TPAMI, 33, 2011. [18] H. Jegou, F. Perronnin, M. Douze, J. Sanchez, P. Perez, and C. Schmid. Aggregating local image descriptors into compact codes. TPAMI, 2012. [19] Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv:1408.5093, 2014. [20] A. Krizhevsky. Learning multiple layers of features from tiny images. Tech Report, 2009. [21] A. Krizhevsky, I. Sutskever, and G. Hinton. Imagenet classiﬁcation with deep convolutional neural networks. In NIPS, 2012. [22] Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural computation, 1989. [23] Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Mu¨ller. Efﬁcient backprop. In Neural Networks: Tricks of the Trade, pages 9–50. Springer, 1998. [24] C.-Y. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu. Deeplysupervised nets. arXiv:1409.5185, 2014. [25] M. Lin, Q. Chen, and S. Yan. Network in network. arXiv:1312.4400, 2013. [26] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dolla´r, and C. L. Zitnick. Microsoft COCO: Common objects in context. In ECCV. 2014. [27] J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015.\n\n[28] G. Montu´far, R. Pascanu, K. Cho, and Y. Bengio. On the number of linear regions of deep neural networks. In NIPS, 2014.\n[29] V. Nair and G. E. Hinton. Rectiﬁed linear units improve restricted boltzmann machines. In ICML, 2010.\n[30] F. Perronnin and C. Dance. Fisher kernels on visual vocabularies for image categorization. In CVPR, 2007.\n[31] T. Raiko, H. Valpola, and Y. LeCun. Deep learning made easier by linear transformations in perceptrons. In AISTATS, 2012.\n[32] S. Ren, K. He, R. Girshick, and J. Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015.\n[33] S. Ren, K. He, R. Girshick, X. Zhang, and J. Sun. Object detection networks on convolutional feature maps. arXiv:1504.06066, 2015.\n[34] B. D. Ripley. Pattern recognition and neural networks. Cambridge university press, 1996.\n[35] A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio. Fitnets: Hints for thin deep nets. In ICLR, 2015.\n[36] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. arXiv:1409.0575, 2014.\n[37] A. M. Saxe, J. L. McClelland, and S. Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv:1312.6120, 2013.\n[38] N. N. Schraudolph. Accelerated gradient descent by factor-centering decomposition. Technical report, 1998.\n[39] N. N. Schraudolph. Centering neural network gradient factors. In Neural Networks: Tricks of the Trade, pages 207–226. Springer, 1998.\n[40] P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. In ICLR, 2014.\n[41] K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.\n[42] R. K. Srivastava, K. Greff, and J. Schmidhuber. Highway networks. arXiv:1505.00387, 2015.\n[43] R. K. Srivastava, K. Greff, and J. Schmidhuber. Training very deep networks. 1507.06228, 2015.\n[44] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015.\n[45] R. Szeliski. Fast surface interpolation using hierarchical basis functions. TPAMI, 1990.\n[46] R. Szeliski. Locally adapted hierarchical basis preconditioning. In SIGGRAPH, 2006.\n[47] T. Vatanen, T. Raiko, H. Valpola, and Y. LeCun. Pushing stochastic gradient towards second-order methods–backpropagation learning with transformations in nonlinearities. In Neural Information Processing, 2013.\n[48] A. Vedaldi and B. Fulkerson. VLFeat: An open and portable library of computer vision algorithms, 2008.\n[49] W. Venables and B. Ripley. Modern applied statistics with s-plus. 1999.\n[50] M. D. Zeiler and R. Fergus. Visualizing and understanding convolutional neural networks. In ECCV, 2014."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "alphago",
      "Paper": "Mastering the game of Go with deep neural networks and tree search",
      "AtlasYear": 2016,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://deepmind-media.storage.googleapis.com/alphago/AlphaGoNaturePaper.pdf",
      "PdfSha256": "9C9184385A3D37B4F4E9D9715270986C43172747B1D08F29093128C1EF878B60",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 926,
          "EndLine": 941,
          "PdfPages": [
            6
          ],
          "Text": "1. Allis, L. V. Searching for Solutions in Games and Artificial Intelligence. PhD thesis, Univ. Limburg, Maastricht, The Netherlands (1994).\n2. van den Herik, H., Uiterwijk, J. W. & van Rijswijck, J. Games solved: now and in the future. Artif. Intell. 134, 277–311 (2002).\n3. Schaeffer, J. The games computers (and people) play. Advances in Computers 52, 189–266 (2000).\n4. Campbell, M., Hoane, A. & Hsu, F. Deep Blue. Artif. Intell. 134, 57–83 (2002). 5. Schaeffer, J. et al. A world championship caliber checkers program. Artif. Intell.\n53, 273–289 (1992). 6. Buro, M. From simple features to sophisticated evaluation functions.\nIn 1st International Conference on Computers and Games, 126–145 (1999). 7. Müller, M. Computer Go. Artif. Intell. 134, 145–179 (2002). 8. Tesauro, G. & Galperin, G. On-line policy improvement using Monte-Carlo\nsearch. In Advances in Neural Information Processing, 1068–1074 (1996). 9. Sheppard, B. World-championship-caliber Scrabble. Artif. Intell. 134, 241–275\n(2002). 10. Bouzy, B. & Helmstetter, B. Monte-Carlo Go developments. In 10th International\nConference on Advances in Computer Games, 159–174 (2003). 11. Coulom, R. Efficient selectivity and backup operators in Monte-Carlo tree\nsearch. In 5th International Conference on Computers and Games, 72–83 (2006). 12. Kocsis, L. & Szepesvári, C. Bandit based Monte-Carlo planning. In 15th European Conference on Machine Learning, 282–293 (2006). 13. Coulom, R. Computing Elo ratings of move patterns in the game of Go. ICGA J. 30, 198–208 (2007). 14. Baudiš, P. & Gailly, J.-L. Pachi: State of the art open source Go program. In Advances in Computer Games, 24–38 (Springer, 2012). 15. Müller, M., Enzenberger, M., Arneson, B. & Segal, R. Fuego – an open-source framework for board games and Go engine based on Monte-Carlo tree search. IEEE Trans. Comput. Intell. AI in Games 2, 259–270 (2010). 16. Gelly, S. & Silver, D. Combining online and offline learning in UCT. In 17th International Conference on Machine Learning, 273–280 (2007).\n\n17. Krizhevsky, A., Sutskever, I. & Hinton, G. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, 1097–1105 (2012).\n18. Lawrence, S., Giles, C. L., Tsoi, A. C. & Back, A. D. Face recognition: a convolutional neural-network approach. IEEE Trans. Neural Netw. 8, 98–113 (1997).\n19. Mnih, V. et al. Human-level control through deep reinforcement learning. Nature 518, 529–533 (2015).\n20. LeCun, Y., Bengio, Y. & Hinton, G. Deep learning. Nature 521, 436–444 (2015). 21. Stern, D., Herbrich, R. & Graepel, T. Bayesian pattern ranking for move\nprediction in the game of Go. In International Conference of Machine Learning, 873–880 (2006). 22. Sutskever, I. & Nair, V. Mimicking Go experts with convolutional neural networks. In International Conference on Artificial Neural Networks, 101–110 (2008). 23. Maddison, C. J., Huang, A., Sutskever, I. & Silver, D. Move evaluation in Go using deep convolutional neural networks. 3rd International Conference on Learning Representations (2015). 24. Clark, C. & Storkey, A. J. Training deep convolutional neural networks to play go. In 32nd International Conference on Machine Learning, 1766–1774 (2015). 25. Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn. 8, 229–256 (1992). 26. Sutton, R., McAllester, D., Singh, S. & Mansour, Y. Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems, 1057–1063 (2000). 27. Sutton, R. & Barto, A. Reinforcement Learning: an Introduction (MIT Press, 1998). 28. Schraudolph, N. N., Dayan, P. & Sejnowski, T. J. Temporal difference learning of position evaluation in the game of Go. Adv. Neural Inf. Process. Syst. 6, 817–824 (1994). 29. Enzenberger, M. Evaluation in Go by a neural network using soft segmentation. In 10th Advances in Computer Games Conference, 97–108 (2003). 267. 30. Silver, D., Sutton, R. & Müller, M. Temporal-difference search in computer Go. Mach. Learn. 87, 183–219 (2012). 31. Levinovitz, A. The mystery of Go, the ancient game that computers still can’t win. Wired Magazine (2014). 32. Mechner, D. All Systems Go. The Sciences 38, 32–37 (1998). 33. Mandziuk, J. Computational intelligence in mind games. In Challenges for Computational Intelligence, 407–442 (2007). 34. Berliner, H. A chronology of computer chess and its literature. Artif. Intell. 10, 201–214 (1978). 35. Browne, C. et al. A survey of Monte-Carlo tree search methods. IEEE Trans. Comput. Intell. AI in Games 4, 1–43 (2012). 36. Gelly, S. et al. The grand challenge of computer Go: Monte Carlo tree search and extensions. Commun. ACM 55, 106–113 (2012). 37. Coulom, R. Whole-history rating: A Bayesian rating system for players of time-varying strength. In International Conference on Computers and Games, 113–124 (2008). 38. KGS. Rating system math. http://www.gokgs.com/help/rmath.html."
        },
        {
          "Section": "Reference section 2",
          "StartLine": 1419,
          "EndLine": 1440,
          "PdfPages": [
            9
          ],
          "Text": "39. Littman, M. L. Markov games as a framework for multi-agent reinforcement learning. In 11th International Conference on Machine Learning, 157–163 (1994).\n40. Knuth, D. E. & Moore, R. W. An analysis of alpha-beta pruning. Artif. Intell. 6, 293–326 (1975).\n41. Sutton, R. Learning to predict by the method of temporal differences. Mach. Learn. 3, 9–44 (1988).\n42. Baxter, J., Tridgell, A. & Weaver, L. Learning to play chess using temporal differences. Mach. Learn. 40, 243–263 (2000).\n43. Veness, J., Silver, D., Blair, A. & Uther, W. Bootstrapping from game tree search. In Advances in Neural Information Processing Systems (2009).\n\n44. Samuel, A. L. Some studies in machine learning using the game of checkers II - recent progress. IBM J. Res. Develop. 11, 601–617 (1967).\n45. Schaeffer, J., Hlynka, M. & Jussila, V. Temporal difference learning applied to a high-performance game-playing program. In 17th International Joint Conference on Artificial Intelligence, 529–534 (2001).\n46. Tesauro, G. TD-gammon, a self-teaching backgammon program, achieves master-level play. Neural Comput. 6, 215–219 (1994).\n47. Dahl, F. Honte, a Go-playing program using neural nets. In Machines that learn to play games, 205–223 (Nova Science, 1999).\n48. Rosin, C. D. Multi-armed bandits with episode context. Ann. Math. Artif. Intell. 61, 203–230 (2011).\n49. Lanctot, M., Winands, M. H. M., Pepels, T. & Sturtevant, N. R. Monte Carlo tree search with heuristic evaluations using implicit minimax backups. In IEEE Conference on Computational Intelligence and Games, 1–8 (2014).\n50. Gelly, S., Wang, Y., Munos, R. & Teytaud, O. Modification of UCT with patterns in Monte-Carlo Go. Tech. Rep. 6062, INRIA (2006).\n51. Silver, D. & Tesauro, G. Monte-Carlo simulation balancing. In 26th International Conference on Machine Learning, 119 (2009).\n52. Huang, S.-C., Coulom, R. & Lin, S.-S. Monte-Carlo simulation balancing in practice. In 7th International Conference on Computers and Games, 81–92 (Springer-Verlag, 2011).\n53. Baier, H. & Drake, P. D. The power of forgetting: improving the last-good-reply policy in Monte Carlo Go. IEEE Trans. Comput. Intell. AI in Games 2, 303–309 (2010).\n54. Huang, S. & Müller, M. Investigating the limits of Monte-Carlo tree search methods in computer Go. In 8th International Conference on Computers and Games, 39–48 (2013).\n55. Segal, R. B. On the scalability of parallel UCT. Computers and Games 6515, 36–47 (2011).\n56. Enzenberger, M. & Müller, M. A lock-free multithreaded Monte-Carlo tree search algorithm. In 12th Advances in Computer Games Conference, 14–20 (2009).\n57. Huang, S.-C., Coulom, R. & Lin, S.-S. Time management for Monte-Carlo tree search applied to the game of Go. In International Conference on Technologies and Applications of Artificial Intelligence, 462–466 (2010).\n58. Gelly, S. & Silver, D. Monte-Carlo tree search and rapid action value estimation in computer Go. Artif. Intell. 175, 1856–1875 (2011).\n59. Baudiš, P. Balancing MCTS by dynamically adjusting the komi value. ICGA J. 34, 131 (2011)."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "tensorflow",
      "Paper": "TensorFlow: A system for large-scale machine learning",
      "AtlasYear": 2016,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/1605.08695",
      "PdfSha256": "42CA0D3F61E4511851688C8B57F3D134A8C64AD9F8AC48AB2B2D930A47234469",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 743,
          "EndLine": 926,
          "PdfPages": [
            13,
            14,
            15,
            16,
            17,
            18
          ],
          "Text": "[1] M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. J. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jo´zefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mane, R. Monga, S. Moore, D. G. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. A. Tucker, V. Vanhoucke, V. Vasudevan, F. B. Vie´gas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng. Tensorﬂow: Large-scale machine learning on heterogeneous distributed systems. CoRR, abs/1603.04467, 2016. arxiv.org/abs/1603.04467. Software available from tensorﬂow.org.\n\n13\n\n\f[2] R. Al-Rfou, G. Alain, A. Almahairi, C. Angermueller, D. Bahdanau, N. Ballas, F. Bastien, J. Bayer, A. Belikov, A. Belopolsky, Y. Bengio, A. Bergeron, J. Bergstra, V. Bisson, J. Bleecher Snyder, N. Bouchard, N. Boulanger-Lewandowski, X. Bouthillier, A. de Bre´bisson, O. Breuleux, P.L. Carrier, K. Cho, J. Chorowski, P. Christiano, T. Cooijmans, M.-A. Coˆte´, M. Coˆte´, A. Courville, Y. N. Dauphin, O. Delalleau, J. Demouth, G. Desjardins, S. Dieleman, L. Dinh, M. Ducoffe, V. Dumoulin, S. Ebrahimi Kahou, D. Erhan, Z. Fan, O. Firat, M. Germain, X. Glorot, I. Goodfellow, M. Graham, C. Gulcehre, P. Hamel, I. Harlouchet, J.-P. Heng, B. Hidasi, S. Honari, A. Jain, S. Jean, K. Jia, M. Korobov, V. Kulkarni, A. Lamb, P. Lamblin, E. Larsen, C. Laurent, S. Lee, S. Lefrancois, S. Lemieux, N. Le´onard, Z. Lin, J. A. Livezey, C. Lorenz, J. Lowin, Q. Ma, P.-A. Manzagol, O. Mastropietro, R. T. McGibbon, R. Memisevic, B. van Merrie¨nboer, V. Michalski, M. Mirza, A. Orlandi, C. Pal, R. Pascanu, M. Pezeshki, C. Raffel, D. Renshaw, M. Rocklin, A. Romero, M. Roth, P. Sadowski, J. Salvatier, F. Savard, J. Schlu¨ter, J. Schulman, G. Schwartz, I. V. Serban, D. Serdyuk, S. Shabanian, E. Simon, S. Spieckermann, S. R. Subramanyam, J. Sygnowski, J. Tanguay, G. van Tulder, J. Turian, S. Urban, P. Vincent, F. Visin, H. de Vries, D. Warde-Farley, D. J. Webb, M. Willson, K. Xu, L. Xue, L. Yao, S. Zhang, and Y. Zhang. Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints, abs/1605.02688, May 2016. arxiv.org/abs/1605.02688.\n[3] A. Angelova, A. Krizhevsky, and V. Vanhoucke. Pedestrian detection with a large-ﬁeld-of-view deep network. In Robotics and Automation (ICRA), 2015 IEEE International Conference on, pages 704–711. IEEE, 2015. CalTech PDF.\n[4] Arvind and D. E. Culler. Annual review of computer science vol. 1, 1986. chapter Dataﬂow Architectures, pages 225–253. 1986. www.dtic.mil/cgi-bin/GetTRDoc?Location=U2& doc=GetTRDoc.pdf&AD=ADA166235.\n[5] J. Ba, V. Mnih, and K. Kavukcuoglu. Multiple object recognition with visual attention. arXiv preprint arXiv:1412.7755, 2014. arxiv.org/abs/1412.7755.\n[6] Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin. A neural probabilistic language model.\n\nJournal of Machine Learning Research, 3:1137– 1155, 2003. www.iro.umontreal.ca/˜lisa/pointeurs/ BengioDucharmeVincentJauvin jmlr.pdf.\n[7] T. Brants and A. Franz. Web 1T 5-gram version 1, 2006. catalog.ldc.upenn.edu/LDC2006T13.\n[8] M. Burrows. The Chubby lock service for looselycoupled distributed systems. In Proceedings of the 7th Symposium on Operating Systems Design and Implementation, OSDI ’06, pages 335–350, Berkeley, CA, USA, 2006. USENIX Association. www.usenix.org/legacy/event/osdi06/tech/full papers/burrows/burrows.pdf.\n[9] R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu. Sample size selection in optimization methods for machine learning. Mathematical Programming, 134(1):127–155, 2012. dx.doi.org/10.1007/s10107012-0572-5.\n[10] C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, and P. Koehn. One billion word benchmark for measuring progress in statistical language modeling. CoRR, abs/1312.3005, 2013. arxiv.org/abs/1312.3005.\n[11] J. Chen, R. Monga, S. Bengio, and R. Jozefowicz. Revisiting distributed synchronous SGD. In International Conference on Learning Representations Workshop Track, 2016. arxiv.org/abs/1604.00981.\n[12] T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang. MXNet: A ﬂexible and efﬁcient machine learning library for heterogeneous distributed systems. In Proceedings of the Workshop on Machine Learning Systems at Neural Information Processing Systems (LearningSys), Dec. 2015. www.cs.cmu.edu/ muli/ﬁle/mxnet-learning-sys.pdf.\n[13] S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer. cuDNN: Efﬁcient primitives for deep learning. arXiv preprint arXiv:1410.0759, 2014. arxiv.org/abs/1410.0759.\n[14] T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman. Project Adam: Building an efﬁcient and scalable deep learning training system. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14), pages 571–582, 2014. www.usenix.org/system/ﬁles/conference/osdi14/ osdi14-paper-chilimbi.pdf.\n\n14\n\n\f[15] S. Chintala.\n\nconvnet-benchmarks, 2016.\n\ngithub.com/soumith/convnet-benchmarks.\n\n[16] E. S. Chung, J. D. Davis, and J. Lee. LINQits: Big data on little clients. In Proceedings of the 40th Annual International Symposium on Computer Architecture, ISCA ’13, pages 261–272, New York, NY, USA, 2013. ACM. doi.acm.org/10.1145/2485922.2485945.\n\n[17] R. Collobert, S. Bengio, and J. Marie´thoz. Torch: A modular machine learning software library. Technical report, IDIAP, 2002. infoscience.epﬂ.ch/record/82802/ﬁles/rr02-46.pdf.\n\n[18] D. Crankshaw, P. Bailis, J. E. Gonzalez, H. Li, Z. Zhang, M. J. Franklin, A. Ghodsi, and M. I. Jordan. The missing piece in complex analytics: Low latency, scalable model management and serving with Velox. In CIDR 2015, Seventh Biennial Conference on Innovative Data Systems Research, Asilomar, CA, USA, January 4-7, 2015, Online Proceedings, 2015. arxiv.org/abs/1409.3809.\n\n[19] H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing. GeePS: Scalable deep learning on distributed GPUs with a GPUspecialized parameter server. In Proceedings of the Eleventh European Conference on Computer Systems, EuroSys ’16, 2016. www.pdl.cmu.edu/PDLFTP/CloudComputing/GeePS-cui-eurosys16.pdf.\n\n[20] A. Dai, C. Olah, and Q. V. Le. Document embedding with paragraph vectors. arXiv preprint arXiv:1507.07998, 2015. arxiv.org/abs/1507.07998.\n\n[21] J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng. Large scale distributed deep networks. In NIPS, 2012. Google Research PDF.\n\n[22] J. Dean and S. Ghemawat. Mapreduce: Simpliﬁed data processing on large clusters. In Proceedings of the 6th Conference on Symposium on Opearting Systems Design & Implementation - Volume 6, OSDI’04, Berkeley, CA, USA, 2004. USENIX Association. research.google.com/archive/mapreduceosdi04.pdf.\n\n[23] A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al. DeVISE: A deep visualsemantic embedding model. In Advances in Neural Information Processing Systems, pages 2121–2129, 2013. research.google.com/pubs/archive/41473.pdf.\n\n[24] J. Gonzalez-Dominguez, I. Lopez-Moreno,\n\nP. J. Moreno, and J. Gonzalez-Rodriguez.\n\nFrame-by-frame language identiﬁcation in\n\nshort utterances using deep neural networks.\n\nNeural Networks, 64:49–58, 2015.\n\nre-\n\nsearch.google.com/en//pubs/archive/42929.pdf.\n\n[25] I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada, pages 2672– 2680, 2014. papers.nips.cc/paper/5423-generativeadversarial-nets.\n\n[26] K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. CoRR, abs/1512.03385, 2015. arxiv.org/abs/1512.03385.\n\n[27] G. Heigold, V. Vanhoucke, A. Senior, P. Nguyen, M. Ranzato, M. Devin, and J. Dean. Multilingual acoustic models using distributed deep neural networks. In Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on, pages 8619–8623. IEEE, 2013. research.google.com/pubs/archive/40807.pdf.\n\n[28] B. Hindman, A. Konwinski, M. Zaharia, A. Ghodsi, A. D. Joseph, R. Katz, S. Shenker, and I. Stoica. Mesos: A platform for ﬁne-grained resource sharing in the data center. In Proceedings of the 8th USENIX Conference on Networked Systems Design and Implementation, NSDI’11, pages 295–308, Berkeley, CA, USA, 2011. USENIX Association. www.cs.berkeley.edu/˜alig/papers/mesos.pdf.\n\n[29] G. E. Hinton. Learning distributed representations of concepts. In Proceedings of the Eighth Annual Conference of the Cognitive Science Society, pages 1–12. Hillsdale, NJ: Erlbaum, 1986. www.cogsci.ucsd.edu/˜ajyu/Teaching/Cogs202 sp13/Readings/hinton86.pdf.\n\n[30] G. E. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Process. Mag., 29(6):82– 97, 2012. www.cs.toronto.edu/˜gdahl/papers/ deepSpeechReviewSPM2012.pdf.\n\n15\n\n\f[31] P. Hunt, M. Konar, F. P. Junqueira, and B. Reed. ZooKeeper: Wait-free coordination for internetscale systems. In Proceedings of the 2010 USENIX Conference on USENIX Annual Technical Conference, USENIXATC’10, pages 11–11, Berkeley, CA, USA, 2010. USENIX Association. www.usenix.org/legacy/event/atc10/tech/full papers/Hunt.pdf.\n[32] S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. CoRR, abs/1502.03167, 2015. arxiv.org/abs/1502.03167.\n[33] B. Jacob et al. gemmlowp: a small selfcontained low-precision GEMM library, 2015. github.com/google/gemmlowp.\n\n[40] A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classiﬁcation with convolutional neural networks. In Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on, pages 1725–1732. IEEE, 2014. research.google.com/pubs/archive/42455.pdf.\n[41] A. Krizhevsky. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997, 2014. arxiv.org/abs/1404.5997.\n[42] A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classiﬁcation with deep convolutional neural networks. In Advances in Neural Information Processing Systems, 2012. papers.nips.cc/paper/4824imagenet-classiﬁcation-with-deep-convolutionalneural-networks.pdf.\n\n[34] B. Jacob, G. Guennebaud, et al. Eigen library for linear algebra. eigen.tuxfamily.org.\n\n[35] S. Jean, K. Cho, R. Memisevic, and Y. Bengio. On using very large target vocabulary for neural machine translation. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1–10, Beijing, China, July 2015. Association for Computational Linguistics. www.aclweb.org/anthology/P15-1001.\n\n[36] Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the ACM International Conference on Multimedia, pages 675–678. ACM, 2014. arxiv.org/pdf/1408.5093.\n\n[37] M. I. Jordan. Serial order: A parallel dis-\n\ntributed processing approach.\n\nICS report\n\n8608, Institute for Cognitive Science, UCSD,\n\nLa Jolla, 1986. cseweb.ucsd.edu/˜gary/PAPER-\n\nSUGGESTIONS/Jordan-TR-8604.pdf.\n\n[38] N. Jouppi.\n\nGoogle supercharges machine\n\nlearning tasks with TPU custom chip, 2016.\n\ncloudplatform.googleblog.com/2016/05/Google-\n\nsupercharges-machine-learning-tasks-with-custom-\n\nchip.html.\n\n[43] H. Larochelle, Y. Bengio, J. Louradour, and P. Lamblin. Exploring strategies for training deep neural networks. Journal of Machine Learning Research, 10:1–40, Jan. 2009. deeplearning.cs.cmu.edu/pdfs/1111/jmlr10 larochelle.pdf.\n[44] A. Lavin and S. Gray. Fast algorithms for convolutional neural networks. CoRR, abs/1509.09308, 2015. arxiv.org/abs/1509.09308.\n[45] Q. Le, M. Ranzato, R. Monga, M. Devin, G. Corrado, K. Chen, J. Dean, and A. Ng. Building highlevel features using large scale unsupervised learning. In ICML’2012, 2012. Google Research PDF.\n[46] M. Li, D. G. Andersen, J. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su. Scaling distributed machine learning with the Parameter Server. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14), pages 583–598, 2014. www.usenix.org/system/ﬁles/conference/osdi14/osdi14paper-chilimbi.pdf.\n[47] M. Li, T. Zhang, Y. Chen, and A. J. Smola. Efﬁcient mini-batch training for stochastic optimization. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’14, pages 661–670, New York, NY, USA, 2014. ACM. www.cs.cmu.edu/˜muli/ﬁle/minibatch sgd.pdf.\n\n[39] R. Jo´zefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu. Exploring the limits of language modeling. CoRR, abs/1602.02410, 2016. arxiv.org/abs/1602.02410.\n\n[48] C. J. Maddison, A. Huang, I. Sutskever, and D. Silver. Move evaluation in Go using deep convolutional neural networks. arXiv preprint arXiv:1412.6564, 2014. arxiv.org/abs/1412.6564.\n\n16\n\n\f[49] F. McSherry, M. Isard, and D. G. Murray. Scalability! But at what COST? In Proceedings of the 15th USENIX Conference on Hot Topics in Operating Systems, HOTOS’15, Berkeley, CA, USA, 2015. USENIX Association. www.usenix.org/system/ﬁles/conference/hotos15/ hotos15-paper-mcsherry.pdf.\n[50] T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efﬁcient estimation of word representations in vector space. In International Conference on Learning Representations: Workshops Track, 2013. arxiv.org/abs/1301.3781.\n[51] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 02 2015. dx.doi.org/10.1038/nature14236.\n\n[58] K. Ovtcharov, O. Ruwase, J.-Y. Kim, J. Fowers, K. Strauss, and E. Chung. Toward accelerating deep learning at scale using specialized logic. In Hot Chips: A Symposium on High Performance Chips. HOTCHIPS, August 2015. research.microsoft.com/apps/pubs/default.aspx?id=246506.\n\n[59] R. Pascanu, T. Mikolov, and Y. Bengio. On the difﬁculty of training recurrent neural networks. In ICML (3), volume 28 of JMLR Proceedings, pages 1310–1318. JMLR.org, 2013. www.jmlr.org/proceedings/papers/v28/pascanu13.pdf.\n\n[60] K. Powell.\n\nNvidia devtech blog post.\n\nblogs.nvidia.com/blog/2015/03/17/digits-devbox/.\n\n[61] J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, and S. Amarasinghe. Halide: A language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. ACM SIGPLAN Notices, 48(6):519– 530, 2013. people.csail.mit.edu/fredo/tmp/Halide5min.pdf.\n\n[52] P. Moritz, R. Nishihara, I. Stoica, and M. I. Jordan. SparkNet: Training deep networks in Spark. In International Conference on Learning Representations, 2016. arxiv.org/abs/1511.06051.\n\n[53] Movidius Ltd. Movidius announces Deep Learning Accelerator and Fathom software framework, 2016. www.movidius.com/news/movidius-announcesdeep-learning-accelerator-and-fathom-softwareframework.\n\n[54] D. G. Murray, F. McSherry, R. Isaacs, M. Isard, P. Barham, and M. Abadi. Naiad: a timely dataﬂow system. In Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles, pages 439–455. ACM, 2013. Microsoft Research PDF.\n\n[55] A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. De Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, et al. Massively parallel methods for deep reinforcement learning. arXiv preprint arXiv:1507.04296, 2015. arxiv.org/abs/1507.04296.\n\n[56] Nervana Systems.\n\nneon,\n\ngithub.com/NervanaSystems/neon.\n\n2016.\n\n[57] NVIDIA Corporation. NCCL: Optimized primitives for collective multi-gpu communication, 2016. github.com/NVIDIA/nccl.\n\n[62] B. Recht, C. Re, S. Wright, and F. Niu. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in Neural Information Processing Systems, pages 693–701, 2011. papers.nips.cc/paper/4390-hogwild-a-lockfree-approach-to-parallelizing-stochastic-gradientdescent.\n[63] C. J. Rossbach, Y. Yu, J. Currey, J.-P. Martin, and D. Fetterly. Dandelion: a compiler and runtime for heterogeneous systems. In Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles, pages 49–68. ACM, 2013. research-srv.microsoft.com/pubs/201110/sosp13dandelion-ﬁnal.pdf.\n[64] D. E. Rumelhart, G. E. Hinton, and R. J. Williams. Learning representations by backpropagating errors. Cognitive modeling, 5:3, 1988. www.cs.toronto.edu/ hinton/absps/naturebp.pdf.\n[65] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. arxiv.org/abs/1409.0575.\n[66] A. Smola and S. Narayanamurthy. An architecture for parallel topic models. Proc.\n\n17\n\n\fVLDB Endow., 3(1-2):703–710, Sept. 2010. vldb.org/pvldb/vldb2010/papers/R63.pdf.\n\nwww.usenix.org/legacy/event/osdi08/tech/full papers/yu y/yu y.pdf.\n\n[67] I. Sutskever, J. Martens, G. E. Dahl, and G. E. Hinton. On the importance of initialization and momentum in deep learning. In Proceedings of the 30th International Conference on Machine Learning (ICML-13), pages 1139–1147. JMLR Workshop and Conference Proceedings, 2013. jmlr.org/proceedings/papers/v28/sutskever13.pdf.\n[68] I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural networks. In NIPS, 2014. papers.nips.cc/paper/5346-sequenceto-sequence-learning-with-neural.\n[69] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR’2015, 2015. arxiv.org/abs/1409.4842.\n\n[75] M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M. J. Franklin, S. Shenker, and I. Stoica. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In Proceedings of the 9th USENIX conference on Networked Systems Design and Implementation. USENIX Association, 2012. www.usenix.org/system/ﬁles/conference/nsdi12/nsdi12ﬁnal138.pdf.\n[76] M. D. Zeiler, M. Ranzato, R. Monga, M. Mao, K. Yang, Q. Le, P. Nguyen, A. Senior, V. Vanhoucke, J. Dean, and G. E. Hinton. On rectiﬁed linear units for speech processing. In ICASSP, 2013. research.google.com/pubs/archive/40811.pdf.\n\n[70] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567, 2015. arxiv.org/abs/1512.00567.\n\n[71] C. tao Chu, S. K. Kim, Y. an Lin, Y. Yu, G. Bradski, K. Olukotun, and A. Y. Ng. Map-reduce for machine learning on multicore. In B. Scho¨lkopf, J. C. Platt, and T. Hoffman, editors, Advances in Neural Information Processing Systems 19, pages 281–288. MIT Press, 2007. papers.nips.cc/paper/3150-mapreduce-for-machine-learning-on-multicore.pdf.\n\n[72] A. Verma, L. Pedrosa, M. Korupolu, D. Oppenheimer, E. Tune, and J. Wilkes. Large-scale cluster management at Google with Borg. In Proceedings of the Tenth European Conference on Computer Systems, page 18. ACM, 2015. research.google.com/pubs/archive/43438.pdf.\n\n[73] O. Vinyals, L. Kaiser, T. Koo, S. Petrov, I. Sutskever, and G. Hinton. Grammar as a foreign language. Technical report, arXiv:1412.7449, 2014. arxiv.org/abs/1412.7449.\n\n[74] Y. Yu, M. Isard, D. Fetterly, M. Budiu, U. Erlingsson, P. K. Gunda, and J. Currey. DryadLINQ: A system for general-purpose distributed dataparallel computing using a high-level language. In Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation, OSDI’08, pages 1–14, Berkeley, CA, USA, 2008. USENIX Association."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "transformer",
      "Paper": "Attention Is All You Need",
      "AtlasYear": 2017,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/1706.03762",
      "PdfSha256": "BDFAA68D8984F0DC02BEACA527B76F207D99B666D31D1DA728EE0728182DF697",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 414,
          "EndLine": 457,
          "PdfPages": [
            10,
            11,
            12
          ],
          "Text": "[1] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.\n[2] Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. CoRR, abs/1409.0473, 2014.\n[3] Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc V. Le. Massive exploration of neural machine translation architectures. CoRR, abs/1703.03906, 2017.\n[4] Jianpeng Cheng, Li Dong, and Mirella Lapata. Long short-term memory-networks for machine reading. arXiv preprint arXiv:1601.06733, 2016.\n10\n\n\f[5] Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. CoRR, abs/1406.1078, 2014.\n[6] Francois Chollet. Xception: Deep learning with depthwise separable convolutions. arXiv preprint arXiv:1610.02357, 2016.\n[7] Junyoung Chung, Çaglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR, abs/1412.3555, 2014.\n[8] Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. Recurrent neural network grammars. In Proc. of NAACL, 2016.\n[9] Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. Convolutional sequence to sequence learning. arXiv preprint arXiv:1705.03122v2, 2017.\n[10] Alex Graves. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013.\n[11] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.\n[12] Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, and Jürgen Schmidhuber. Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001.\n[13] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.\n[14] Zhongqiang Huang and Mary Harper. Self-training PCFG grammars with latent annotations across languages. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 832–841. ACL, August 2009.\n[15] Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410, 2016.\n[16] Łukasz Kaiser and Samy Bengio. Can active memory replace attention? In Advances in Neural Information Processing Systems, (NIPS), 2016.\n[17] Łukasz Kaiser and Ilya Sutskever. Neural GPUs learn algorithms. In International Conference on Learning Representations (ICLR), 2016.\n[18] Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu. Neural machine translation in linear time. arXiv preprint arXiv:1610.10099v2, 2017.\n[19] Yoon Kim, Carl Denton, Luong Hoang, and Alexander M. Rush. Structured attention networks. In International Conference on Learning Representations, 2017.\n[20] Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.\n[21] Oleksii Kuchaiev and Boris Ginsburg. Factorization tricks for LSTM networks. arXiv preprint arXiv:1703.10722, 2017.\n[22] Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130, 2017.\n[23] Minh-Thang Luong, Quoc V. Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser. Multi-task sequence to sequence learning. arXiv preprint arXiv:1511.06114, 2015.\n[24] Minh-Thang Luong, Hieu Pham, and Christopher D Manning. Effective approaches to attentionbased neural machine translation. arXiv preprint arXiv:1508.04025, 2015.\n11\n\n\f[25] Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. Building a large annotated corpus of english: The penn treebank. Computational linguistics, 19(2):313–330, 1993.\n[26] David McClosky, Eugene Charniak, and Mark Johnson. Effective self-training for parsing. In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 152–159. ACL, June 2006.\n[27] Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. A decomposable attention model. In Empirical Methods in Natural Language Processing, 2016.\n[28] Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017.\n[29] Slav Petrov, Leon Barrett, Romain Thibaux, and Dan Klein. Learning accurate, compact, and interpretable tree annotation. In Proceedings of the 21st International Conference on Computational Linguistics and 44th Annual Meeting of the ACL, pages 433–440. ACL, July 2006.\n[30] Ofir Press and Lior Wolf. Using the output embedding to improve language models. arXiv preprint arXiv:1608.05859, 2016.\n[31] Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015.\n[32] Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.\n[33] Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929–1958, 2014.\n[34] Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. End-to-end memory networks. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems 28, pages 2440–2448. Curran Associates, Inc., 2015.\n[35] Ilya Sutskever, Oriol Vinyals, and Quoc VV Le. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, pages 3104–3112, 2014.\n[36] Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567, 2015.\n[37] Vinyals & Kaiser, Koo, Petrov, Sutskever, and Hinton. Grammar as a foreign language. In Advances in Neural Information Processing Systems, 2015.\n[38] Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.\n[39] Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu. Deep recurrent models with fast-forward connections for neural machine translation. CoRR, abs/1606.04199, 2016.\n[40] Muhua Zhu, Yue Zhang, Wenliang Chen, Min Zhang, and Jingbo Zhu. Fast and accurate shift-reduce constituent parsing. In Proceedings of the 51st Annual Meeting of the ACL (Volume 1: Long Papers), pages 434–443. ACL, August 2013."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "bert",
      "Paper": "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding",
      "AtlasYear": 2018,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/1810.04805",
      "PdfSha256": "5692A5514787A8C6727B4FF3B726A3385798BC68E12138D1D4AF83947E2ACF6E",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 559,
          "EndLine": 619,
          "PdfPages": [
            10,
            11,
            12
          ],
          "Text": "Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual string embeddings for sequence labeling. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1638–1649.\nRami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2018. Character-level language modeling with deeper self-attention. arXiv preprint arXiv:1808.04444.\nRie Kubota Ando and Tong Zhang. 2005. A framework for learning predictive structures from multiple tasks and unlabeled data. Journal of Machine Learning Research, 6(Nov):1817–1853.\nLuisa Bentivogli, Bernardo Magnini, Ido Dagan, Hoa Trang Dang, and Danilo Giampiccolo. 2009. The ﬁfth PASCAL recognizing textual entailment challenge. In TAC. NIST.\nJohn Blitzer, Ryan McDonald, and Fernando Pereira. 2006. Domain adaptation with structural correspondence learning. In Proceedings of the 2006 conference on empirical methods in natural language processing, pages 120–128. Association for Computational Linguistics.\nSamuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In EMNLP. Association for Computational Linguistics.\nPeter F Brown, Peter V Desouza, Robert L Mercer, Vincent J Della Pietra, and Jenifer C Lai. 1992. Class-based n-gram models of natural language. Computational linguistics, 18(4):467–479.\nDaniel Cer, Mona Diab, Eneko Agirre, Inigo LopezGazpio, and Lucia Specia. 2017. Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, Vancouver, Canada. Association for Computational Linguistics.\nCiprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005.\nZ. Chen, H. Zhang, X. Zhang, and L. Zhao. 2018. Quora question pairs.\nChristopher Clark and Matt Gardner. 2018. Simple and effective multi-paragraph reading comprehension. In ACL.\n\nKevin Clark, Minh-Thang Luong, Christopher D Manning, and Quoc Le. 2018. Semi-supervised sequence modeling with cross-view training. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1914– 1925.\nRonan Collobert and Jason Weston. 2008. A uniﬁed architecture for natural language processing: Deep neural networks with multitask learning. In Proceedings of the 25th international conference on Machine learning, pages 160–167. ACM.\nAlexis Conneau, Douwe Kiela, Holger Schwenk, Lo¨ıc Barrault, and Antoine Bordes. 2017. Supervised learning of universal sentence representations from natural language inference data. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 670–680, Copenhagen, Denmark. Association for Computational Linguistics.\nAndrew M Dai and Quoc V Le. 2015. Semi-supervised sequence learning. In Advances in neural information processing systems, pages 3079–3087.\nJ. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. FeiFei. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09.\nWilliam B Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005).\nWilliam Fedus, Ian Goodfellow, and Andrew M Dai. 2018. Maskgan: Better text generation via ﬁlling in the . arXiv preprint arXiv:1801.07736.\nDan Hendrycks and Kevin Gimpel. 2016. Bridging nonlinearities and stochastic regularizers with gaussian error linear units. CoRR, abs/1606.08415.\nFelix Hill, Kyunghyun Cho, and Anna Korhonen. 2016. Learning distributed representations of sentences from unlabelled data. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics.\nJeremy Howard and Sebastian Ruder. 2018. Universal language model ﬁne-tuning for text classiﬁcation. In ACL. Association for Computational Linguistics.\nMinghao Hu, Yuxing Peng, Zhen Huang, Xipeng Qiu, Furu Wei, and Ming Zhou. 2018. Reinforced mnemonic reader for machine reading comprehension. In IJCAI.\nYacine Jernite, Samuel R. Bowman, and David Sontag. 2017. Discourse-based objectives for fast unsupervised sentence representation learning. CoRR, abs/1705.00557.\n\n\fMandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In ACL.\nRyan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-thought vectors. In Advances in neural information processing systems, pages 3294–3302.\nQuoc Le and Tomas Mikolov. 2014. Distributed representations of sentences and documents. In International Conference on Machine Learning, pages 1188–1196.\nHector J Levesque, Ernest Davis, and Leora Morgenstern. 2011. The winograd schema challenge. In Aaai spring symposium: Logical formalizations of commonsense reasoning, volume 46, page 47.\nLajanugen Logeswaran and Honglak Lee. 2018. An efﬁcient framework for learning sentence representations. In International Conference on Learning Representations.\nBryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017. Learned in translation: Contextualized word vectors. In NIPS.\nOren Melamud, Jacob Goldberger, and Ido Dagan. 2016. context2vec: Learning generic context embedding with bidirectional LSTM. In CoNLL.\nTomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26, pages 3111–3119. Curran Associates, Inc.\nAndriy Mnih and Geoffrey E Hinton. 2009. A scalable hierarchical distributed language model. In D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou, editors, Advances in Neural Information Processing Systems 21, pages 1081–1088. Curran Associates, Inc.\nAnkur P Parikh, Oscar Ta¨ckstro¨m, Dipanjan Das, and Jakob Uszkoreit. 2016. A decomposable attention model for natural language inference. In EMNLP.\nJeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In Empirical Methods in Natural Language Processing (EMNLP), pages 1532– 1543.\nMatthew Peters, Waleed Ammar, Chandra Bhagavatula, and Russell Power. 2017. Semi-supervised sequence tagging with bidirectional language models. In ACL.\nMatthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018a. Deep contextualized word representations. In NAACL.\n\nMatthew Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018b. Dissecting contextual word embeddings: Architecture and representation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1499–1509.\nAlec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding with unsupervised learning. Technical report, OpenAI.\nPranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392.\nMinjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2017. Bidirectional attention ﬂow for machine comprehension. In ICLR.\nRichard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.\nFu Sun, Linyang Li, Xipeng Qiu, and Yang Liu. 2018. U-net: Machine reading comprehension with unanswerable questions. arXiv preprint arXiv:1810.06638.\nWilson L Taylor. 1953. Cloze procedure: A new tool for measuring readability. Journalism Bulletin, 30(4):415–433.\nErik F Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the conll-2003 shared task: Language-independent named entity recognition. In CoNLL.\nJoseph Turian, Lev Ratinov, and Yoshua Bengio. 2010. Word representations: A simple and general method for semi-supervised learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, ACL ’10, pages 384–394.\nAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, pages 6000–6010.\nPascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. 2008. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, pages 1096–1103. ACM.\nAlex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018a. Glue: A multi-task benchmark and analysis platform\n\n\ffor natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355.\nWei Wang, Ming Yan, and Chen Wu. 2018b. Multigranularity hierarchical attention fusion networks for reading comprehension and question answering. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics.\nAlex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2018. Neural network acceptability judgments. arXiv preprint arXiv:1805.12471.\nAdina Williams, Nikita Nangia, and Samuel R Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In NAACL.\nYonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144.\nJason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014. How transferable are features in deep neural networks? In Advances in neural information processing systems, pages 3320–3328.\nAdams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V Le. 2018. QANet: Combining local convolution with global self-attention for reading comprehension. In ICLR.\nRowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. Swag: A large-scale adversarial dataset for grounded commonsense inference. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP).\nYukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pages 19–27."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "gpt-2",
      "Paper": "Language Models are Unsupervised Multitask Learners",
      "AtlasYear": 2019,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf",
      "PdfSha256": "D9D852E2894556E73F53CB22B7C605A9643D6F0B19BF604B429ED6192FA24F4E",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 272,
          "EndLine": 379,
          "PdfPages": [
            10,
            11,
            12
          ],
          "Text": "Al-Rfou, R., Choe, D., Constant, N., Guo, M., and Jones, L. Character-level language modeling with deeper self-attention. arXiv preprint arXiv:1808.04444, 2018.\nAlberti, C., Lee, K., and Collins, M. A bert baseline for the natural questions. arXiv preprint arXiv:1901.08634, 2019.\nAlcorn, M. A., Li, Q., Gong, Z., Wang, C., Mai, L., Ku, W.-S., and Nguyen, A. Strike (with) a pose: Neural networks are easily fooled by strange poses of familiar objects. arXiv preprint arXiv:1811.11553, 2018.\nAmodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., Chen, G., et al. Deep speech 2: End-to-end speech recognition in english and mandarin. In International Conference on Machine Learning, pp. 173–182, 2016.\nArtetxe, M., Labaka, G., Agirre, E., and Cho, K. Unsupervised neural machine translation. arXiv preprint arXiv:1710.11041, 2017.\nArtetxe, M., Labaka, G., and Agirre, E. An effective approach to unsupervised machine translation. arXiv preprint arXiv:1902.01313, 2019.\n5Preliminary code for downloading and using the small model is available at https://github.com/openai/gpt-2\n\nBa, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.\n\nBajgar, O., Kadlec, R., and Kleindienst, J. Embracing data abundance: Booktest dataset for reading comprehension. arXiv preprint arXiv:1610.00956, 2016.\n\nBarz, B. and Denzler, J. Do we train on test data? purging cifar of near-duplicates. arXiv preprint arXiv:1902.00423, 2019.\n\nBengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. A neural probabilistic language model. Journal of machine learning research, 3(Feb):1137–1155, 2003.\n\nBowman, S. R., Pavlick, E., Grave, E., Van Durme, B., Wang, A., Hula, J., Xia, P., Pappagari, R., McCoy, R. T., Patel, R., et al. Looking for elmo’s friends: Sentence-level pretraining beyond language modeling. arXiv preprint arXiv:1812.10860, 2018.\n\nCaruana, R. Multitask learning. Machine learning, 28(1):41–75, 1997.\n\nChelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005, 2013.\n\nCollobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., and Kuksa, P. Natural language processing (almost) from scratch. Journal of Machine Learning Research, 12(Aug):2493– 2537, 2011.\n\nConneau, A., Kiela, D., Schwenk, H., Barrault, L., and Bordes, A. Supervised learning of universal sentence representations from natural language inference data. arXiv preprint arXiv:1705.02364, 2017a.\n\nConneau, A., Lample, G., Ranzato, M., Denoyer, L., and Je´gou, H. Word translation without parallel data. arXiv preprint arXiv:1710.04087, 2017b.\n\nDai, A. M. and Le, Q. V. Semi-supervised sequence learning. In Advances in neural information processing systems, pp. 3079– 3087, 2015.\n\nDai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R. Transformer-xl: Attentive language models beyond a ﬁxed-length context. arXiv preprint arXiv:1901.02860, 2019.\n\nDavies, M.\n\nThe 14 billion word iweb corpus.\n\nhttps://corpus.byu.edu/iWeb/, 2018.\n\nDehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł. Universal transformers. arXiv preprint arXiv:1807.03819, 2018.\n\nDevlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pretraining of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.\n\nDinan, E., Roller, S., Shuster, K., Fan, A., Auli, M., and Weston, J. Wizard of wikipedia: Knowledge-powered conversational agents. arXiv preprint arXiv:1811.01241, 2018.\n\nFan, A., Lewis, M., and Dauphin, Y. Hierarchical neural story generation. arXiv preprint arXiv:1805.04833, 2018.\n\n\fLanguage Models are Unsupervised Multitask Learners\n\nFinn, C., Abbeel, P., and Levine, S. Model-agnostic metalearning for fast adaptation of deep networks. arXiv preprint arXiv:1703.03400, 2017.\nGehrmann, S., Deng, Y., and Rush, A. M. Bottom-up abstractive summarization. arXiv preprint arXiv:1808.10792, 2018.\nGillick, D., Brunk, C., Vinyals, O., and Subramanya, A. Multilingual language processing from bytes. arXiv preprint arXiv:1512.00103, 2015.\nGong, C., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y. Frage: frequency-agnostic word representation. In Advances in Neural Information Processing Systems, pp. 1341–1352, 2018.\nGrave, E., Joulin, A., and Usunier, N. Improving neural language models with a continuous cache. arXiv preprint arXiv:1612.04426, 2016.\nHe, K., Zhang, X., Ren, S., and Sun, J. Identity mappings in deep residual networks. In European conference on computer vision, pp. 630–645. Springer, 2016.\nHestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Ali, M., Yang, Y., and Zhou, Y. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409, 2017.\nHill, F., Bordes, A., Chopra, S., and Weston, J. The goldilocks principle: Reading children’s books with explicit memory representations. arXiv preprint arXiv:1511.02301, 2015.\nHill, F., Cho, K., and Korhonen, A. Learning distributed representations of sentences from unlabelled data. arXiv preprint arXiv:1602.03483, 2016.\nHoang, L., Wiseman, S., and Rush, A. M. Entity tracking improves cloze-style reading comprehension. arXiv preprint arXiv:1810.02891, 2018.\nHoward, J. and Ruder, S. Universal language model ﬁne-tuning for text classiﬁcation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), volume 1, pp. 328–339, 2018.\nJelinek, F. and Mercer, R. L. Interpolated estimation of markov source parameters from sparse data. In Proceedings of the Workshop on Pattern Recognition in Practice, Amsterdam, The Netherlands: North-Holland, May., 1980.\nJia, R. and Liang, P. Adversarial examples for evaluating reading comprehension systems. arXiv preprint arXiv:1707.07328, 2017.\nJozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410, 2016.\nKaiser, L., Gomez, A. N., Shazeer, N., Vaswani, A., Parmar, N., Jones, L., and Uszkoreit, J. One model to learn them all. arXiv preprint arXiv:1706.05137, 2017.\nKarpathy, A., Johnson, J., and Fei-Fei, L. Visualizing and understanding recurrent networks. arXiv preprint arXiv:1506.02078, 2015.\n\nKirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., GrabskaBarwinska, A., et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, pp. 201611835, 2017.\nKiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S. Skip-thought vectors. In Advances in neural information processing systems, pp. 3294–3302, 2015.\nKrizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classiﬁcation with deep convolutional neural networks. In Advances in neural information processing systems, pp. 1097–1105, 2012.\nKwiatkowski, T., Palomaki, J., Rhinehart, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., et al. Natural questions: a benchmark for question answering research. 2019.\nLake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. Building machines that learn and think like people. Behavioral and Brain Sciences, 40, 2017.\nLample, G., Conneau, A., Denoyer, L., and Ranzato, M. Unsupervised machine translation using monolingual corpora only. arXiv preprint arXiv:1711.00043, 2017.\nLevesque, H., Davis, E., and Morgenstern, L. The winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning, 2012.\nLevy, O. and Goldberg, Y. Neural word embedding as implicit matrix factorization. In Advances in neural information processing systems, pp. 2177–2185, 2014.\nLiu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N. Generating wikipedia by summarizing long sequences. arXiv preprint arXiv:1801.10198, 2018.\nMcCann, B., Bradbury, J., Xiong, C., and Socher, R. Learned in translation: Contextualized word vectors. In Advances in Neural Information Processing Systems, pp. 6294–6305, 2017.\nMcCann, B., Keskar, N. S., Xiong, C., and Socher, R. The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730, 2018.\nMerity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.\nMikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pp. 3111–3119, 2013.\nNallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al. Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023, 2016.\nPaperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Ferna´ndez, R. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.\nPennington, J., Socher, R., and Manning, C. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp. 1532–1543, 2014.\n\n\fLanguage Models are Unsupervised Multitask Learners\n\nPeters, M. E. and Lecocq, D. Content extraction using diverse feature sets. In Proceedings of the 22nd International Conference on World Wide Web, pp. 89–90. ACM, 2013.\nPeters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. Deep contextualized word representations. arXiv preprint arXiv:1802.05365, 2018.\nRadford, A., Jozefowicz, R., and Sutskever, I. Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444, 2017.\nRadford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. 2018.\nRamachandran, P., Liu, P. J., and Le, Q. V. Unsupervised pretraining for sequence to sequence learning. arXiv preprint arXiv:1611.02683, 2016.\nRecht, B., Roelofs, R., Schmidt, L., and Shankar, V. Do cifar-10 classiﬁers generalize to cifar-10? arXiv preprint arXiv:1806.00451, 2018.\nReddy, S., Chen, D., and Manning, C. D. Coqa: A conversational question answering challenge. arXiv preprint arXiv:1808.07042, 2018.\nSchwartz, R., Sap, M., Konstas, I., Zilles, L., Choi, Y., and Smith, N. A. Story cloze task: Uw nlp system. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pp. 52–55, 2017.\nSee, A., Liu, P. J., and Manning, C. D. Get to the point: Summarization with pointer-generator networks. arXiv preprint arXiv:1704.04368, 2017.\nSennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015.\nSubramanian, S., Trischler, A., Bengio, Y., and Pal, C. J. Learning general purpose distributed sentence representations via large scale multi-task learning. arXiv preprint arXiv:1804.00079, 2018.\nSutskever, I., Vinyals, O., and Le, Q. V. Sequence to sequence learning with neural networks. In Advances in neural information processing systems, pp. 3104–3112, 2014.\nSutskever, I., Jozefowicz, R., Gregor, K., Rezende, D., Lillicrap, T., and Vinyals, O. Towards principled unsupervised learning. arXiv preprint arXiv:1511.06440, 2015.\nTrichelair, P., Emami, A., Cheung, J. C. K., Trischler, A., Suleman, K., and Diaz, F. On the evaluation of common-sense reasoning in natural language understanding. arXiv preprint arXiv:1811.01778, 2018.\nTrinh, T. H. and Le, Q. V. A simple method for commonsense reasoning. arXiv preprint arXiv:1806.02847, 2018.\nVaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems, pp. 5998–6008, 2017.\nVinyals, O. and Le, Q. A neural conversational model. arXiv preprint arXiv:1506.05869, 2015.\n\nVinyals, O., Fortunato, M., and Jaitly, N. Pointer networks. In Advances in Neural Information Processing Systems, pp. 2692– 2700, 2015.\nWang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.\nWeston, J. E. Dialog-based language learning. In Advances in Neural Information Processing Systems, pp. 829–837, 2016.\nWieting, J. and Kiela, D. No training required: Exploring random encoders for sentence classiﬁcation. arXiv preprint arXiv:1901.10444, 2019.\nWolf, T., Sanh, V., Chaumond, J., and Delangue, C. Transfertransfo: A transfer learning approach for neural network based conversational agents. arXiv preprint arXiv:1901.08149, 2019.\nYogatama, D., d’Autume, C. d. M., Connor, J., Kocisky, T., Chrzanowski, M., Kong, L., Lazaridou, A., Ling, W., Yu, L., Dyer, C., et al. Learning and evaluating general linguistic intelligence. arXiv preprint arXiv:1901.11373, 2019."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "pytorch",
      "Paper": "PyTorch: An Imperative Style, High-Performance Deep Learning Library",
      "AtlasYear": 2019,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/1912.01703",
      "PdfSha256": "B47FB010F6872701D6D3BDD8270497BE700C9961E2EC2CAC2B08A154A43D81F2",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 186,
          "EndLine": 292,
          "PdfPages": [
            9,
            10,
            11,
            12
          ],
          "Text": "[1] Yangqing \"Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor\" Darrell. \"caffe: Convolutional architecture for fast feature embedding\". \"arXiv preprint arXiv:1408.5093\", \"2014\".\n[2] Frank Seide and Amit Agarwal. Cntk: Microsoft’s open-source deep-learning toolkit. In Proceedings of the 22Nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pages 2135–2135, New York, NY, USA, 2016. ACM.\n9\n\n\f[3] Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. TensorFlow: Largescale machine learning on heterogeneous systems, 2015. Software available from tensorﬂow.org.\n\n[4] Theano Development Team. Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints, abs/1605.02688, May 2016.\n\n[5] Seiya Tokui, Kenta Oono, Shohei Hido, and Justin Clayton. Chainer: a next-generation open source framework for deep learning. In Proceedings of Workshop on Machine Learning Systems (LearningSys) in The Twenty-ninth Annual Conference on Neural Information Processing Systems (NIPS), 2015.\n\n[6] Ronan Collobert, Samy Bengio, and Johnny Mariéthoz. Torch: a modular machine learning software library. Technical report, Idiap, 2002.\n\n[7] G. Neubig, C. Dyer, Y. Goldberg, A. Matthews, W. Ammar, A. Anastasopoulos, M. Ballesteros, D. Chiang, D. Clothiaux, T. Cohn, K. Duh, M. Faruqui, C. Gan, D. Garrette, Y. Ji, L. Kong, A. Kuncoro, G. Kumar, C. Malaviya, P. Michel, Y. Oda, M. Richardson, N. Saphra, S. Swayamdipta, and P. Yin. DyNet: The Dynamic Neural Network Toolkit. ArXiv e-prints, January 2017.\n\n[8] Philip S. Abrams. An APL Machine. PhD thesis, Stanford University, 1970.\n\n[9] The MathWorks, Inc., Natick, Massachusetts, United States. MATLAB and Statistics Toolbox.\n\n[10] R Core Team. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria.\n\n[11] Jeff Bezanson, Alan Edelman, Stefan Karpinski, and Viral B Shah. Julia: A fresh approach to numerical computing. SIAM review, 59(1):65–98, 2017.\n\n[12] Travis Oliphant. NumPy: A guide to NumPy. http://www.numpy.org/.\n\nUSA: Trelgol Publishing, 2006.\n\n[13] Gaël Guennebaud, Benoît Jacob, et al. Eigen v3. http://eigen.tuxfamily.org, 2010.\n\n[14] Y LeCun and L Bottou. Lush reference manual. Technical report, code available at http://lush.sourceforge.net, 2002.\n\n[15] Atilim Gunes Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind. Automatic differentiation in machine learning: A survey. J. Mach. Learn. Res., 18(1):5595–5637, January 2017.\n\n[16] Dougal Maclaurin. Modeling, Inference and Optimization with Composable Differentiable Procedures. PhD thesis, Harvard University, April 2016.\n\n[17] Matthew Johnson et. al. Jax. https://github.com/google/jax, 2018.\n\n[18] Mike Innes et. al. Flux.jl. https://github.com/FluxML/Flux.jl, 2018.\n\n[19] Eric Jones, Travis Oliphant, Pearu Peterson, et al. SciPy: Open source scientiﬁc tools for Python, 2001–. http://www.scipy.org/.\n\n[20] Wes McKinney. Data structures for statistical computing in python. In Proceedings of the 9th Python in Science Conference, 51-56, 2010.\n\n[21] Pierre Sermanet, Koray Kavukcuoglu, and Yann LeCun. Eblearn: Open-source energy-based learning in c++. In 2009 21st IEEE International Conference on Tools with Artiﬁcial Intelligence, pages 693–697. IEEE, 2009.\n\n10\n\n\f[22] Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan D. Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. cudnn: Efﬁcient primitives for deep learning. CoRR, abs/1410.0759, 2014.\n\n[23] Andrew Lavin. maxdnn: An efﬁcient convolution kernel for deep learning with maxwell gpus, January 2015.\n\n[24] Andrew Lavin and Scott Gray. Fast algorithms for convolutional neural networks. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4013–4021, 2016.\n\n[25] Ronan Collobert, Koray Kavukcuoglu, and Clément Farabet. Torch7: A matlab-like environment for machine learning. In NIPS 2011, 2011.\n\n[26] Richard Gabriel. The rise of worse is better. http://dreamsongs.com/RiseOfWorseIsBetter.html.\n\n[27] Yann LeCun and Corinna Cortes. http://yann.lecun.com/exdb/mnist/.\n\nMNIST handwritten digit database.\n\n[28] Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, John Quan, Stephen Gaffney, Stig Petersen, Karen Simonyan, Tom Schaul, Hado van Hasselt, David Silver, Timothy P. Lillicrap, Kevin Calderone, Paul Keet, Anthony Brunasso, David Lawrence, Anders Ekermo, Jacob Repp, and Rodney Tsing. Starcraft II: A new challenge for reinforcement learning. CoRR, abs/1708.04782, 2017.\n\n[29] DMLC. Dlpack: Open in memory tensor structure. https://github.com/dmlc/dlpack.\n\n[30] Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. In NIPS Workshop, 2017.\n\n[31] Dan Piponi. Automatic differentiation, C++ templates, and photogrammetry. J. Graphics, GPU, & Game Tools, 9(4):41–55, 2004.\n\n[32] Holger Leuck and Hans-Hellmut Nagel. Automatic differentiation facilitates of-integration into steering-angle-based road vehicle tracking. In 1999 Conference on Computer Vision and Pattern Recognition (CVPR ’99), 23-25 June 1999, Ft. Collins, CO, USA, pages 2360–2365, 1999.\n\n[33] The Python team.\n\nThe cpython\n\nhttps://wiki.python.org/moin/GlobalInterpreterLock.\n\nglobal\n\ninterpreter\n\nlock.\n\n[34] Giovanni Petrantoni and Jörg Wollenschläger. Nimtorch. https://github.com/fragcolorxyz/nimtorch.\n\n[35] Austin Huang, Junji Hashimoto, https://github.com/hasktorch/hasktorch.\n\nand Sam Stites.\n\nHasktorch.\n\n[36] G. Synnaeve, Z. Lin, J. Gehring, D. Gant, V. Mella, V. Khalidov, N. Carion, and N. Usunier. Forward modeling for partial observation strategy games - a starcraft defogger. In Advances in Neural Information Processing Systems, pages 10761–10771, 2018.\n\n[37] The PyTorch team. Torch Script. https://pytorch.org/docs/stable/jit.html.\n\n[38] Justin Luitjens. Cuda streams. GPU technology conference, 2014.\n\n[39] Emery D. Berger, Kathryn S. McKinley, Robert D. Blumofe, and Paul R. Wilson. Hoard: A scalable memory allocator for multithreaded applications. In Proceedings of the Ninth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS IX, pages 117–128, New York, NY, USA, 2000. ACM.\n\n[40] J. Evans. A scalable concurrent malloc(3) implementation for freebsd. In In BSDCan — The Technical BSD Conference, May 2006.\n\n[41] S. Ghemawat and P. Menage. Tcmalloc: Thread-caching malloc.\n\n11\n\n\f[42] Benjamin Recht, Christopher Ré, Stephen J. Wright, and Feng Niu. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12-14 December 2011, Granada, Spain., pages 693–701, 2011.\n[43] Matthew Hertz and Emery D. Berger. Quantifying the performance of garbage collection vs. explicit memory management. In Proceedings of the 20th Annual ACM SIGPLAN Conference on Object-oriented Programming, Systems, Languages, and Applications, OOPSLA ’05, pages 313–326, New York, NY, USA, 2005. ACM.\n[44] The PyTorch team. Pytorch Autograd Proﬁler. https://pytorch.org/docs/1.0.1/autograd.html#proﬁler."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "scaling-laws",
      "Paper": "Scaling Laws for Neural Language Models",
      "AtlasYear": 2020,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2001.08361",
      "PdfSha256": "A41BD7877FD1A6BCBBA096B2619618BD2F90E02488F2365644903EB2F7C6A494",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 2006,
          "EndLine": 2146,
          "PdfPages": [
            28,
            29,
            30,
            31
          ],
          "Text": "[ACDE12] Eduardo G Altmann, Giampaolo Cristadoro, and Mirko Degli Esposti. On the origin of longrange correlations in texts. Proceedings of the National Academy of Sciences, 109(29):11582– 11587, 2012. 25\n\n[AS17]\n\nMadhu S. Advani and Andrew M. Saxe. High-dimensional dynamics of generalization error in neural networks. arXiv, 2017, 1710.03667. 11, 18, 22\n\n[BB01]\n\nMichele Banko and Eric Brill. Scaling to very very large corpora for natural language disambiguation. In Proceedings of the 39th annual meeting on association for computational linguistics, pages 26–33. Association for Computational Linguistics, 2001. 18\n\n[BHMM18] Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine learning and the bias-variance trade-off. arXiv, 2018, 1812.11118. 18\n\n[Bia12]\n\nGÃŠrard Biau. Analysis of a random forests model. Journal of Machine Learning Research, 13(Apr):1063–1095, 2012. 18\n\n[CGRS19] Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. CoRR, abs/1904.10509, 2019, 1904.10509. URL http://arxiv.org/ abs/1904.10509. 19\n\n[DCLT18] [DGV+18]\n\nJacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding, 2018, arXiv:1810.04805. 2\nMostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers. CoRR, abs/1807.03819, 2018, 1807.03819. URL http://arxiv.org/ abs/1807.03819. 6, 9, 23, 24\n\n[EP94]\n\nWerner Ebeling and Thorsten Pöschel. Entropy and long-range correlations in literary english. EPL (Europhysics Letters), 26(4):241, 1994. 25\n\n[Fou]\n\nThe Common Crawl Foundation. Common crawl. URL http://commoncrawl.org. 7\n\n[GARD18] [GJS+19]\n\nGuy Gur-Ari, Daniel A. Roberts, and Ethan Dyer. Gradient descent happens in a tiny subspace. 2018, arXiv:1812.04754. 18\nMario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart. Scaling description of generalization with number of parameters in deep learning. arXiv, 2019, 1901.01608. 18\n\n[GKX19]\n\nBehrooz Ghorbani, Shankar Krishnan, and Ying Xiao. An investigation into neural net optimization via hessian eigenvalue density. CoRR, abs/1901.10159, 2019, 1901.10159. URL http://arxiv.org/abs/1901.10159. 18\n\n[Goo01]\n\nJoshua Goodman. A bit of progress in language modeling. CoRR, cs.CL/0108005, 2001. URL http://arxiv.org/abs/cs.CL/0108005. 18\n\n[GRK17] Scott Gray, Alec Radford, and Diederik P Kingma. Gpu kernels for block-sparse weights. openai.com, 2017. 19\n\n[HAD19]\n\nJoel Hestness, Newsha Ardalani, and Gregory Diamos. Beyond human-level accuracy: Computational challenges in deep learning. In Proceedings of the 24th Symposium on Principles and Practice of Parallel Programming, PPoPP ’19, pages 1–14, New York, NY, USA, 2019. ACM. doi:10.1145/3293883.3295710. 18\n\n28\n\n\f[HCC+18] Yanping Huang, Yonglong Cheng, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, and Zhifeng Chen. Gpipe: Efﬁcient training of giant neural networks using pipeline parallelism. CoRR, abs/1811.06965, 2018, 1811.06965. URL http://arxiv.org/abs/1811.06965. 19\n\n[HNA+17] Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically, 2017, 1712.00409. 18\n\n[JGH18]\n\nArthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in neural information processing systems, pages 8571–8580, 2018. 18\n\n[KB14]\n\nDiederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2014, 1412.6980. 7\n\n[Kom19] Aran Komatsuzaki. One epoch is all you need, 2019, arXiv:1906.06669. 18\n\n[KSH12]\n\nAlex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classiﬁcation with deep convolutional neural networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 1, NIPS’12, pages 1097–1105, USA, 2012. Curran Associates Inc. URL http://dl.acm.org/citation.cfm?id=2999134.2999257. 19\n\n[LCG+19] Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations, 2019, 1909.11942. 9\n\n[LOG+19]\n\nYinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019, 1907.11692. URL http://arxiv.org/abs/ 1907.11692. 2\n\n[LSP+18]\n\nPeter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating wikipedia by summarizing long sequences. arXiv:1801.10198 [cs], 2018, 1801.10198. URL http://arxiv.org/abs/1801.10198. 2, 6\n\n[LT16]\n\nHenry W Lin and Max Tegmark. Criticality in formal languages and statistical physics. arXiv preprint arXiv:1606.06737, 2016. 25\n\n[LXS+19]\n\nJaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha SohlDickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent, 2019, arXiv:1902.06720. 18\n\n[MKAT18] Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. An empirical model of large-batch training, 2018, arXiv:1812.06162. 3, 5, 6, 12, 13, 21\n\n[Pap18]\n\nVardan Papyan. The full spectrum of deep net hessians at scale: Dynamics with sample size. CoRR, abs/1811.07062, 2018, 1811.07062. URL http://arxiv.org/abs/1811.07062. 18\n\n[RNSS18]\n\nAlec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. URL https://s3-us-west-2. amazonaws. com/openaiassets/research-covers/languageunsupervised/language understanding paper. pdf, 2018. 2, 6\n\n[RRBS19a] Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. A constructive prediction of the generalization error across scales, 2019, 1909.12673. 18\n\n[RRBS19b] Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. A constructive prediction of the generalization error across scales, 2019, arXiv:1909.12673. 18\n\n[RSR+19]\n\nColin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a uniﬁed text-to-text transformer, 2019, arXiv:1910.10683. 2\n\n[RWC+19] Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. openai.com, 2019. 2, 5, 6, 7, 8\n\n[SCP+18]\n\nNoam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake Hechtman. Mesh-tensorﬂow: Deep learning for supercomputers, 2018, 1811.02084. 19\n\n[SHB15]\n\nRico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. CoRR, 2015, 1508.07909. 6\n\n29\n\n\f[SLA+18] [SS18] [THK18] [TL19] [VSP+17]\n[VWB16] [Was06] [WPN+19] [WRH17] [WYL19] [YDY+19] [ZK16] [ZKZ+15]\n[ZLN+19]\n\nChristopher J. Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl. Measuring the effects of data parallelism on neural network training, 2018, arXiv:1811.03600. 12\nNoam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. CoRR, abs/1804.04235, 2018, 1804.04235. URL http://arxiv.org/abs/1804.04235. 7\nStefan Thurner, Rudolf Hanel, and Peter Klimek. Introduction to the theory of complex systems. Oxford University Press, 2018. 18\nMingxing Tan and Quoc V. Le. Efﬁcientnet: Rethinking model scaling for convolutional neural networks. CoRR, abs/1905.11946, 2019, 1905.11946. URL http://arxiv.org/abs/1905. 11946. 18\nAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 5998–6008. Curran Associates, Inc., 2017. URL http://papers.nips.cc/paper/7181-attention-is-all-you-need.pdf. 2, 6\nAndreas Veit, Michael Wilber, and Serge Belongie. Residual networks behave like ensembles of relatively shallow networks, 2016, arXiv:1605.06431. 8, 18\nLarry Wasserman. All of nonparametric statistics. Springer Science & Business Media, 2006. 18\nAlex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems, 2019, 1905.00537. 2\nYu-Xiong Wang, Deva Ramanan, and Martial Hebert. Growing a brain: Fine-tuning by increasing model capacity. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul 2017. doi:10.1109/cvpr.2017.323. 19\nWei Wen, Feng Yan, and Hai Li. Autogrow: Automatic layer growing in deep convolutional networks, 2019, 1906.02909. 19\nZhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. Xlnet: Generalized autoregressive pretraining for language understanding, 2019, arXiv:1906.08237. 2\nSergey Zagoruyko and Nikos Komodakis. Wide residual networks. Procedings of the British Machine Vision Conference 2016, 2016. doi:10.5244/c.30.87. 18\nYukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. 2015 IEEE International Conference on Computer Vision (ICCV), Dec 2015. doi:10.1109/iccv.2015.11. 7\nGuodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George E. Dahl, Christopher J. Shallue, and Roger B. Grosse. Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model. CoRR, abs/1907.04164, 2019, 1907.04164. URL http://arxiv.org/abs/1907.04164. 12, 18\n\n30"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "gpt-3",
      "Paper": "Language Models are Few-Shot Learners",
      "AtlasYear": 2020,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2005.14165",
      "PdfSha256": "97FD272F1FDFC18677462D0292F5FBF26CA86B4D1B485C2DBA03269B643A0E83",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 2545,
          "EndLine": 2705,
          "PdfPages": [
            68,
            69,
            70,
            71,
            72,
            73,
            74,
            75,
            76
          ],
          "Text": "[ADG+16] Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas. Learning to learn by gradient descent by gradient descent. In Advances in neural information processing systems, pages 3981–3989, 2016.\n[AI19] WeChat AI. Tr-mt (ensemble), December 2019.\n[AJF19] Roee Aharoni, Melvin Johnson, and Orhan Firat. Massively multilingual neural machine translation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019.\n[BBDIW20] Su Lin Blodgett, Solon Barocas, Hal Daume´ III, and Hanna Wallach. Language (technology) is power: A critical survey of “bias” in nlp. arXiv preprint arXiv:2005.14050, 2020.\n[BCFL13] Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1533–1544, 2013.\n[BDD+09] Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini. The ﬁfth PASCAL recognizing textual entailment challenge. 2009.\n[BES10] Stefano Baccianella, Andrea Esuli, and Fabrizio Sebastiani. Sentiwordnet 3.0: an enhanced lexical resource for sentiment analysis and opinion mining. In Lrec, volume 10, pages 2200–2204, 2010.\n[BHDD+06] Roy Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. The second PASCAL recognising textual entailment challenge. 2006.\n[BHT+20] Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et al. Experience grounds language. arXiv preprint arXiv:2004.10151, 2020.\n[BLC13] Yoshua Bengio, Nicholas Le´onard, and Aaron C. Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. Arxiv, 2013.\n[BZB+19] Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language. arXiv preprint arXiv:1911.11641, 2019.\n[Car97] Rich Caruana. Multitask learning. Machine learning, 28(1), 1997.\n[CB78] Susan Carey and Elsa Bartlett. Acquiring a single new word. Proceedings of the Stanford Child Language Conference, 1978.\n[CCE+18] Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv, abs/1803.05457, 2018.\n[CGRS19] Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers, 2019.\n[CHI+18] Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. Quac : Question answering in context. Arxiv, 2018.\n[CLC+19] Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difﬁculty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.\n[CLY+19] Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Learning universal image-text representations. arXiv preprint arXiv:1909.11740, 2019.\n[Cra17] Kate Crawford. The trouble with bias. NIPS 2017 Keynote, 2017.\n[DCLT18] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.\n68\n\n\f[DGM06] Ido Dagan, Oren Glickman, and Bernardo Magnini. The PASCAL recognising textual entailment challenge. In Machine learning challenges. evaluating predictive uncertainty, visual object classiﬁcation, and recognising textual entailment, pages 177–190. Springer, 2006.\n[DGV+18] Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers. Arxiv, 2018.\n[DHKH14] Nadir Durrani, Barry Haddow, Philipp Koehn, and Kenneth Heaﬁeld. Edinburgh’s phrase-based machine translation systems for wmt-14. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 97–104, 2014.\n[DL15] Andrew M. Dai and Quoc V. Le. Semi-supervised sequence learning. In Advances in neural information processing systems, 2015.\n[DMST19] Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. The CommitmentBank: Investigating projection in naturally occurring discourse. 2019. To appear in proceedings of Sinn und Bedeutung 23. Data can be found at https://github.com/mcdm/CommitmentBank/.\n[DSC+16] Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel. Rl2: Fast reinforcement learning via slow reinforcement learning. ArXiv, abs/1611.02779, 2016.\n[DWD+19] Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. arXiv preprint arXiv:1903.00161, 2019.\n[DYY+19] Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a ﬁxed-length context. Arxiv, 2019.\n[EOAG18] Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. Understanding back-translation at scale. arXiv preprint arXiv:1808.09381, 2018.\n[FAL17] Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. ArXiv, abs/1703.03400, 2017.\n[Fyo00] Yaroslav Fyodorov. A natural logic inference system, 2000.\n[GG19] Hila Gonen and Yoav Goldberg. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. arXiv preprint arXiv:1903.03862, 2019.\n[GLT+20] Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. Realm: Retrievalaugmented language model pre-training. arXiv preprint arXiv:2002.08909, 2020.\n[GMDD07] Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. The third PASCAL recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL workshop on textual entailment and paraphrasing, pages 1–9. Association for Computational Linguistics, 2007.\n[Gra16] Alex Graves. Adaptive computation time for recurrent neural networks. Arxiv, 2016.\n[GSL+18] Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R Bowman, and Noah A Smith. Annotation artifacts in natural language inference data. arXiv preprint arXiv:1803.02324, 2018.\n[GSR19] Sebastian Gehrmann, Hendrik Strobelt, and Alexander M. Rush. Gltr: Statistical detection and visualization of generated text. arXiv preprint arXiv: 1906.04043, 2019.\n[GWC+18] Jiatao Gu, Yong Wang, Yun Chen, Kyunghyun Cho, and Victor OK Li. Meta-learning for low-resource neural machine translation. arXiv preprint arXiv:1808.08437, 2018.\n[HB20] Daniel Hernandez and Tom Brown. Ai and efﬁciency, May 2020.\n[HBFC19] Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. CoRR, abs/1904.09751, 2019.\n[HLW+20] Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. Pretrained transformers improve out of distribution robustness. arXiv preprint arXiv:2004.06100, 2020.\n69\n\n\f[HNA+17] Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409, 2017.\n[HR18] Jeremy Howard and Sebastian Ruder. Universal language model ﬁne-tuning for text classiﬁcation. arXiv preprint arXiv:1801.06146, 2018.\n[HVD15] Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.\n[HYC01] Sepp Hochreiter, A Steven Younger, and Peter R Conwell. Learning to Learn Using Gradient Descent. In International Conference on Artiﬁcial Neural Networks, pages 87–94. Springer, 2001.\n[HZJ+19] Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. Reducing sentiment bias in language models via counterfactual evaluation. arXiv preprint arXiv:1911.03064, 2019.\n[IBGC+14] Mohit Iyyer, Jordan Boyd-Graber, Leonardo Claudino, Richard Socher, and Hal Daume´ III. A neural network for factoid question answering over paragraphs. In Empirical Methods in Natural Language Processing, 2014.\n[IDCBE19] Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. Automatic detection of generated text is easiest when humans are fooled. arXiv preprint arXiv:1911.00650, 2019.\n[JCWZ17] Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017.\n[JN20] Zheng Junyuan and Gamma Lab NYC. Numeric transformer - albert, March 2020.\n[JVS+16] Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410, 2016.\n[JYS+19] Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. TinyBERT: Distilling BERT for natural language understanding. arXiv preprint arXiv:1909.10351, 2019.\n[JZC+19] Ying Ju, Fubang Zhao, Shijie Chen, Bowen Zheng, Xuefeng Yang, and Yunfeng Liu. Technical report on conversational question answering. arXiv preprint arXiv:1909.10772, 2019.\n[KCR+18] Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of North American Chapter of the Association for Computational Linguistics (NAACL), 2018.\n[KKS+20] Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. Uniﬁedqa: Crossing format boundaries with a single qa system. arXiv preprint arXiv:2005.00700, 2020.\n[KMB20] Sarah E. Kreps, Miles McCain, and Miles Brundage. All the news that’s ﬁt to fabricate: Ai-generated text as a tool of media misinformation, 2020.\n[KMH+20] Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models, 2020.\n[KPR+19] Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redﬁeld, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics, 2019.\n[KR16] Yoon Kim and Alexander M. Rush. Sequence-level knowledge distillation. Arxiv, 2016.\n[LB02] Edward Loper and Steven Bird. Nltk: The natural language toolkit, 2002.\n[LC19] Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291, 2019.\n70\n\n\f[LCG+19] Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. ALBERT: A lite BERT for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019.\n[LCH+20] Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. Adversarial training for large neural language models. arXiv preprint arXiv:2004.08994, 2020.\n[LDL19] Zhongyang Li, Xiao Ding, and Ting Liu. Story ending prediction by transferable bert. arXiv preprint arXiv:1905.07504, 2019.\n[LDM12] Hector Levesque, Ernest Davis, and Leora Morgenstern. The Winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning, 2012.\n[LGG+20] Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. Multilingual denoising pre-training for neural machine translation. arXiv preprint arXiv:2001.08210, 2020.\n[LGH+15] Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, and Ye-Yi Wang. Representation learning using multi-task deep neural networks for semantic classiﬁcation and information retrieval. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2015.\n[LH17] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.\n[LHCG19a] Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. Improving multi-task deep neural networks via knowledge distillation for natural language understanding. arXiv preprint arXiv:1904.09482, 2019.\n[LHCG19b] Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. Multi-task deep neural networks for natural language understanding. arXiv preprint arXiv:1901.11504, 2019.\n[Lin20] Tal Linzen. How can we accelerate progress towards human-like linguistic generalization? arXiv preprint arXiv:2005.00955, 2020.\n[LLG+19] Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019.\n[LM17] Ke Li and Jitendra Malik. Learning to optimize neural nets. arXiv preprint arXiv:1703.00441, 2017.\n[LOG+19] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019.\n[LPP+20] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Ku¨ttler, Mike Lewis, Wen-tau Yih, Tim Rockta¨schel, Sebastian Riedel, and Kiela Douwe. Retrieval-augmented generation for knowledge-intensive nlp tasks. arXiv preprint arXiv:2005.11401, 2020.\n[LSP+18] Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating Wikipedia by summarizing long sequences. arXiv preprint arXiv:1801.10198, 2018.\n[LWS+20] Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E. Gonzalez. Train large, then compress: Rethinking model size for efﬁcient training and inference of transformers, 2020.\n[LXL+17] Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683, 2017.\n[LYN+20] Sheng-Chieh Lin, Jheng-Hong Yang, Rodrigo Nogueira, Ming-Feng Tsai, Chuan-Ju Wang, and Jimmy Lin. Tttttackling winogrande schemas. arXiv preprint arXiv:2003.08380, 2020.\n[Mac92] David. MacKay. Information-based objective functions for active data selection. Neural Computation, 1992.\n71\n\n\f[MBXS17] Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. Learned in translation: Contextualized word vectors. In Advances in Neural Information Processing Systems, pages 6294–6305, 2017.\n[MCCD13] Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efﬁcient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013.\n[MCH+16] Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. A corpus and evaluation framework for deeper understanding of commonsense stories. arXiv preprint arXiv:1604.01696, 2016.\n[MCKS18] Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. ArXiv, abs/1809.02789, 2018.\n[MKAT18] Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. An empirical model of large-batch training, 2018.\n[MKM+94] Mitchell Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert MacIntyre, Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger. The penn treebank: annotating predicate argument structure. In Proceedings of the workshop on Human Language Technology, pages 114–119. Association for Computational Linguistics, 1994.\n[MKXS18] Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730, 2018.\n[MPL19] R Thomas McCoy, Ellie Pavlick, and Tal Linzen. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. arXiv preprint arXiv:1902.01007, 2019.\n[MWZ+18] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting, 2018.\n[NBR20] Moin Nadeem, Anna Bethke, and Siva Reddy. Stereoset: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456, 2020.\n[NK19] Timothy Niven and Hung-Yu Kao. Probing neural network comprehension of natural language arguments. arXiv preprint arXiv:1907.07355, 2019.\n[Nor09] Peter Norvig. Natural language corpus data, 2009.\n[NvNvdG19] Malvina Nissim, Rik van Noord, and Rob van der Goot. Fair is better than sensational: Man is to doctor as woman is to doctor. arXiv preprint arXiv:1905.09866, 2019.\n[NWD+19] Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. Adversarial nli: A new benchmark for natural language understanding. arXiv preprint arXiv:1910.14599, 2019.\n[oR16] University of Regensburg. Fascha, 2016.\n[PCC18] Mohammad Taher Pilehvar and Jose Camacho-Collados. WIC: 10,000 example pairs for evaluating context-sensitive representations. arXiv preprint arXiv:1808.09121, 2018.\n[PFB18] Jason Phang, Thibault Fe´vry, and Samuel R. Bowman. Sentence encoders on STILTs: Supplementary training on intermediate labeled-data tasks. arXiv preprint arXiv:1811.01088, 2018.\n[PHR+18] Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. Collecting diverse natural language inference problems for sentence representation evaluation. In Proceedings of EMNLP, 2018.\n[PKL+16] Denis Paperno, Germa´n Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Ferna´ndez. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.\n[PNZtY18] Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen tau Yih. Dissecting contextual word embeddings: Architecture and representation, 2018.\n[Pos18] Matt Post. A call for clarity in reporting BLEU scores. arXiv preprint arXiv:1804.08771, 2018.\n72\n\n\f[PSM14] Jeffrey Pennington, Richard Socher, and Christopher Manning. GloVe: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014.\n[QIA20] QIANXIN. Sa-net on albert (ensemble), April 2020.\n[QMZH19] Yusu Qian, Urwa Muaz, Ben Zhang, and Jae Won Hyun. Reducing gender bias in word-level language models with a gender-equalizing loss function. arXiv preprint arXiv:1905.12801, 2019.\n[RBG11] Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI Spring Symposium Series, 2011.\n[RCM19] Siva Reddy, Danqi Chen, and Christopher D Manning. Coqa: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266, 2019.\n[RCP+17] Scott Reed, Yutian Chen, Thomas Paine, Aa¨ron van den Oord, SM Eslami, Danilo Rezende, Oriol Vinyals, and Nando de Freitas. Few-shot autoregressive density estimation: Towards learning to learn distributions. arXiv preprint arXiv:1710.10304, 2017.\n[RJL18] Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for squad. arXiv preprint arXiv:1806.03822, 2018.\n[RL16] Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. ICLR 2017 (oral), 2016.\n[RLL+19] Qiu Ran, Yankai Lin, Peng Li, Jie Zhou, and Zhiyuan Liu. NumNet: Machine reading comprehension with numerical reasoning. In Proceedings of EMNLP, 2019.\n[RNLVD18] Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. Gender bias in coreference resolution. arXiv preprint arXiv:1804.09301, 2018.\n[RNSS18] Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018.\n[Ros12] R.S. Ross. Guide for conducting risk assessments. NIST Special Publication, 2012.\n[RRBS19] Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. A constructive prediction of the generalization error across scales, 2019.\n[RRS20] Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020.\n[RSR+19] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a uniﬁed text-to-text transformer, 2019.\n[RWC+19] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners, 2019.\n[SBBC19] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale, 2019.\n[SBC+19] Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGufﬁe, and Jasmine Wang. Release strategies and the social impacts of language models, 2019.\n[SCNP19] Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. The woman worked as a babysitter: On biases in language generation. arXiv preprint arXiv:1909.01326, 2019.\n[SDCW19] Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.\n[SDSE19] Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. Green AI. CoRR, abs/1907.10597, 2019.\n[SHB15] Rico Sennrich, Barry Haddow, and Alexandra Birch. Improving neural machine translation models with monolingual data. arXiv preprint arXiv:1511.06709, 2015.\n73\n\n\f[SMM+17] Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.\n[SPP+19] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism, 2019.\n[SS20] Timo Schick and Hinrich Schu¨tze. Exploiting cloze questions for few-shot text classiﬁcation and natural language inference. arXiv preprint arXiv:2001.07676, 2020.\n[STQ+19] Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. MASS: Masked sequence to sequence pre-training for language generation. arXiv preprint arXiv:1905.02450, 2019.\n[TFR+17] Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23–30. IEEE, 2017.\n[TL05] Peter D. Turney and Michael L. Littman. Corpus-based learning of analogies and semantic relations. CoRR, abs/cs/0508103, 2005.\n[TL18] Trieu H. Trinh and Quoc V. Le. A simple method for commonsense reasoning. arXiv preprint arXiv:1806.02847, 2018.\n[TLBS03] Peter D. Turney, Michael L. Littman, Jeffrey Bigham, and Victor Shnayder. Combining independent modules to solve multiple-choice synonym and analogy problems. CoRR, cs.CL/0309035, 2003.\n[Tur20] Project Turing. Microsoft research blog, Feb 2020.\n[VBL+16] Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching Networks for One Shot Learning. In Advances in neural information processing systems, pages 3630–3638, 2016.\n[VSP+17] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, 2017.\n[WPN+19] Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems, pages 3261–3275, 2019.\n[WXH+18] Yiren Wang, Yingce Xia, Tianyu He, Fei Tian, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu. Multi-agent dual learning. ICLR 2019, 2018.\n[XDH+19] Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V. Le. Unsupervised data augmentation for consistency training, 2019.\n[YdC+19] Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, et al. Learning and evaluating general linguistic intelligence. arXiv preprint arXiv:1901.11373, 2019.\n[YDY+19] Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. XLNet: Generalized autoregressive pretraining for language understanding. arXiv preprint arXiv:1906.08237, 2019.\n[ZHB+19] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really ﬁnish your sentence? arXiv preprint arXiv:1905.07830, 2019.\n[ZHR+19] Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. Defending against neural fake news. arXiv preprint arXiv:1905.12616, 2019.\n[ZLL+18] Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. ReCoRD: Bridging the gap between human and machine commonsense reading comprehension. arXiv preprint arXiv:1810.12885, 2018.\n[ZSW+19a] Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences, 2019.\n74\n\n\f[ZSW+19b] Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. ArXiv, abs/1909.08593, 2019.\n75"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "ddpm",
      "Paper": "Denoising Diffusion Probabilistic Models",
      "AtlasYear": 2020,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2006.11239",
      "PdfSha256": "AEE5E07A802E8DFD2A386374C94FD61D1D056CB7E1E0FEC4F28E8120FF5D8505",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 591,
          "EndLine": 669,
          "PdfPages": [
            9,
            10,
            11,
            12
          ],
          "Text": "[1] Guillaume Alain, Yoshua Bengio, Li Yao, Jason Yosinski, Eric Thibodeau-Laufer, Saizheng Zhang, and Pascal Vincent. GSNs: generative stochastic networks. Information and Inference: A Journal of the IMA, 5(2):210–249, 2016.\n[2] Florian Bordes, Sina Honari, and Pascal Vincent. Learning to generate samples from noise through infusion training. In International Conference on Learning Representations, 2017.\n[3] Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high ﬁdelity natural image synthesis. In International Conference on Learning Representations, 2019.\n[4] Tong Che, Ruixiang Zhang, Jascha Sohl-Dickstein, Hugo Larochelle, Liam Paull, Yuan Cao, and Yoshua Bengio. Your GAN is secretly an energy-based model and you should use discriminator driven latent sampling. arXiv preprint arXiv:2003.06060, 2020.\n[5] Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, pages 6571–6583, 2018.\n[6] Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel. PixelSNAIL: An improved autoregressive generative model. In International Conference on Machine Learning, pages 863–871, 2018.\n[7] Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.\n9\n\n\f[8] Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio Ranzato. Residual energy-based models for text generation. arXiv preprint arXiv:2004.11714, 2020.\n[9] Laurent Dinh, David Krueger, and Yoshua Bengio. NICE: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516, 2014.\n[10] Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using Real NVP. arXiv preprint arXiv:1605.08803, 2016.\n[11] Yilun Du and Igor Mordatch. Implicit generation and modeling with energy based models. In Advances in Neural Information Processing Systems, pages 3603–3613, 2019.\n[12] Ruiqi Gao, Yang Lu, Junpei Zhou, Song-Chun Zhu, and Ying Nian Wu. Learning generative ConvNets via multi-grid modeling and sampling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9155–9164, 2018.\n[13] Ruiqi Gao, Erik Nijkamp, Diederik P Kingma, Zhen Xu, Andrew M Dai, and Ying Nian Wu. Flow contrastive estimation of energy-based models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7518–7528, 2020.\n[14] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, pages 2672–2680, 2014.\n[15] Anirudh Goyal, Nan Rosemary Ke, Surya Ganguli, and Yoshua Bengio. Variational walkback: Learning a transition operator as a stochastic recurrent net. In Advances in Neural Information Processing Systems, pages 4392–4402, 2017.\n[16] Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt, and David Duvenaud. FFJORD: Free-form continuous dynamics for scalable reversible generative models. In International Conference on Learning Representations, 2019.\n[17] Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky. Your classiﬁer is secretly an energy based model and you should treat it like one. In International Conference on Learning Representations, 2020.\n[18] Karol Gregor, Frederic Besse, Danilo Jimenez Rezende, Ivo Danihelka, and Daan Wierstra. Towards conceptual compression. In Advances In Neural Information Processing Systems, pages 3549–3557, 2016.\n[19] Prahladh Harsha, Rahul Jain, David McAllester, and Jaikumar Radhakrishnan. The communication complexity of correlation. In Twenty-Second Annual IEEE Conference on Computational Complexity (CCC’07), pages 10–23. IEEE, 2007.\n[20] Marton Havasi, Robert Peharz, and José Miguel Hernández-Lobato. Minimal random code learning: Getting bits back from compressed model parameters. In International Conference on Learning Representations, 2019.\n[21] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems, pages 6626–6637, 2017.\n[22] Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-VAE: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations, 2017.\n[23] Jonathan Ho, Xi Chen, Aravind Srinivas, Yan Duan, and Pieter Abbeel. Flow++: Improving ﬂow-based generative models with variational dequantization and architecture design. In International Conference on Machine Learning, 2019.\n[24] Sicong Huang, Alireza Makhzani, Yanshuai Cao, and Roger Grosse. Evaluating lossy compression rates of deep generative models. In International Conference on Machine Learning, 2020.\n[25] Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu. Video pixel networks. In International Conference on Machine Learning, pages 1771–1779, 2017.\n[26] Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu. Efﬁcient neural audio synthesis. In International Conference on Machine Learning, pages 2410–2419, 2018.\n[27] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. In International Conference on Learning Representations, 2018.\n[28] Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages\n10\n\n\f4401–4410, 2019.\n[29] Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Training generative adversarial networks with limited data. arXiv preprint arXiv:2006.06676v1, 2020.\n[30] Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of StyleGAN. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8110–8119, 2020.\n[31] Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.\n[32] Diederik P Kingma and Prafulla Dhariwal. Glow: Generative ﬂow with invertible 1x1 convolutions. In Advances in Neural Information Processing Systems, pages 10215–10224, 2018.\n[33] Diederik P Kingma and Max Welling. Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114, 2013.\n[34] Diederik P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. Improved variational inference with inverse autoregressive ﬂow. In Advances in Neural Information Processing Systems, pages 4743–4751, 2016.\n[35] John Lawson, George Tucker, Bo Dai, and Rajesh Ranganath. Energy-inspired models: Learning with sampler-induced distributions. In Advances in Neural Information Processing Systems, pages 8501–8513, 2019.\n[36] Daniel Levy, Matt D. Hoffman, and Jascha Sohl-Dickstein. Generalizing Hamiltonian Monte Carlo with neural networks. In International Conference on Learning Representations, 2018.\n[37] Lars Maaløe, Marco Fraccaro, Valentin Liévin, and Ole Winther. BIVA: A very deep hierarchy of latent variables for generative modeling. In Advances in Neural Information Processing Systems, pages 6548–6558, 2019.\n[38] Jacob Menick and Nal Kalchbrenner. Generating high ﬁdelity images with subscale pixel networks and multidimensional upscaling. In International Conference on Learning Representations, 2019.\n[39] Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations, 2018.\n[40] Alex Nichol. VQ-DRAW: A sequential discrete VAE. arXiv preprint arXiv:2003.01599, 2020.\n[41] Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu. On the anatomy of MCMC-based maximum likelihood learning of energy-based models. arXiv preprint arXiv:1903.12370, 2019.\n[42] Erik Nijkamp, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu. Learning non-convergent non-persistent short-run MCMC toward energy-based model. In Advances in Neural Information Processing Systems, pages 5233–5243, 2019.\n[43] Georg Ostrovski, Will Dabney, and Remi Munos. Autoregressive quantile networks for generative modeling. In International Conference on Machine Learning, pages 3936–3945, 2018.\n[44] Ryan Prenger, Rafael Valle, and Bryan Catanzaro. WaveGlow: A ﬂow-based generative network for speech synthesis. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3617–3621. IEEE, 2019.\n[45] Ali Razavi, Aaron van den Oord, and Oriol Vinyals. Generating diverse high-ﬁdelity images with VQVAE-2. In Advances in Neural Information Processing Systems, pages 14837–14847, 2019.\n[46] Danilo Rezende and Shakir Mohamed. Variational inference with normalizing ﬂows. In International Conference on Machine Learning, pages 1530–1538, 2015.\n[47] Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In International Conference on Machine Learning, pages 1278–1286, 2014.\n[48] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 234–241. Springer, 2015.\n[49] Tim Salimans and Durk P Kingma. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. In Advances in Neural Information Processing Systems, pages 901–909, 2016.\n[50] Tim Salimans, Diederik Kingma, and Max Welling. Markov Chain Monte Carlo and variational inference: Bridging the gap. In International Conference on Machine Learning, pages 1218–1226, 2015.\n11\n\n\f[51] Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In Advances in Neural Information Processing Systems, pages 2234–2242, 2016.\n[52] Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P Kingma. PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modiﬁcations. In International Conference on Learning Representations, 2017.\n[53] Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265, 2015.\n[54] Jiaming Song, Shengjia Zhao, and Stefano Ermon. A-NICE-MC: Adversarial training for MCMC. In Advances in Neural Information Processing Systems, pages 5140–5150, 2017.\n[55] Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems, pages 11895–11907, 2019.\n[56] Yang Song and Stefano Ermon. Improved techniques for training score-based generative models. arXiv preprint arXiv:2006.09011, 2020.\n[57] Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.\n[58] Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. International Conference on Machine Learning, 2016.\n[59] Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu. Conditional image generation with PixelCNN decoders. In Advances in Neural Information Processing Systems, pages 4790–4798, 2016.\n[60] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008, 2017.\n[61] Pascal Vincent. A connection between score matching and denoising autoencoders. Neural Computation, 23(7):1661–1674, 2011.\n[62] Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. Cnn-generated images are surprisingly easy to spot...for now. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.\n[63] Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7794–7803, 2018.\n[64] Auke J Wiggers and Emiel Hoogeboom. Predictive sampling with forecasting autoregressive models. arXiv preprint arXiv:2002.09928, 2020.\n[65] Hao Wu, Jonas Köhler, and Frank Noé. Stochastic normalizing ﬂows. arXiv preprint arXiv:2002.06707, 2020.\n[66] Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European Conference on Computer Vision (ECCV), pages 3–19, 2018.\n[67] Jianwen Xie, Yang Lu, Song-Chun Zhu, and Yingnian Wu. A theory of generative convnet. In International Conference on Machine Learning, pages 2635–2644, 2016.\n[68] Jianwen Xie, Song-Chun Zhu, and Ying Nian Wu. Synthesizing dynamic patterns by spatial-temporal generative convnet. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7093–7101, 2017.\n[69] Jianwen Xie, Zilong Zheng, Ruiqi Gao, Wenguan Wang, Song-Chun Zhu, and Ying Nian Wu. Learning descriptor networks for 3d shape synthesis and analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8629–8638, 2018.\n[70] Jianwen Xie, Song-Chun Zhu, and Ying Nian Wu. Learning energy-based spatial-temporal generative convnets for dynamic patterns. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019.\n[71] Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. LSUN: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015.\n[72] Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "clip",
      "Paper": "Learning Transferable Visual Models From Natural Language Supervision",
      "AtlasYear": 2021,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2103.00020",
      "PdfSha256": "6478B6E571A7D6FCD846D8EF77BFD60C285F1986ABB8F475EEDC43DE403074F5",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 1286,
          "EndLine": 1289,
          "PdfPages": [
            27
          ],
          "Text": "Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. Tensorﬂow: A system for large-scale machine learning. In 12th {USENIX} symposium on operating systems design and implementation ({OSDI} 16), pp. 265–283, 2016.\nAlayrac, J.-B., Recasens, A., Schneider, R., Arandjelovic´, R., Ramapuram, J., De Fauw, J., Smaira, L., Dieleman, S., and Zisserman, A. Self-supervised multimodal versatile networks. arXiv preprint arXiv:2006.16228, 2020.\nAlcorn, M. A., Li, Q., Gong, Z., Wang, C., Mai, L., Ku, W.S., and Nguyen, A. Strike (with) a pose: Neural networks are easily fooled by strange poses of familiar objects. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4845–4854, 2019.\nAndreas, J., Klein, D., and Levine, S. Learning with latent language. arXiv preprint arXiv:1711.00482, 2017."
        },
        {
          "Section": "Reference section 2",
          "StartLine": 1294,
          "EndLine": 1296,
          "PdfPages": [
            27
          ],
          "Text": "Assiri, Y. Stochastic optimization of plain convolutional neural networks with simple methods. arXiv preprint arXiv:2001.08856, 2020.\nBachman, P., Hjelm, R. D., and Buchwalter, W. Learning representations by maximizing mutual information across views. In Advances in Neural Information Processing Systems, pp. 15535–15545, 2019.\nBarbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Advances in Neural Information Processing Systems, pp. 9453–9463, 2019."
        },
        {
          "Section": "Reference section 3",
          "StartLine": 1300,
          "EndLine": 1601,
          "PdfPages": [
            27,
            28,
            29,
            30,
            31,
            32,
            33,
            34,
            35,
            36
          ],
          "Text": "Barnard, K., Duygulu, P., Forsyth, D., Freitas, N. d., Blei, D. M., and Jordan, M. I. Matching words and pictures. Journal of machine learning research, 3(Feb):1107–1135, 2003.\nBechmann, A. and Bowker, G. C. Unsupervised by any other name: Hidden layers of knowledge production in artiﬁcial intelligence on social media. Big Data & Society, 6(1):205395171881956, January 2019. doi: 10.1177/ 2053951718819569. URL https://doi.org/10. 1177/2053951718819569.\nBengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. A neural probabilistic language model. Journal of machine learning research, 3(Feb):1137–1155, 2003.\nBhargava, S. and Forsyth, D. Exposing and correcting the gender bias in image captioning datasets and models. arXiv preprint arXiv:1912.00578, 2019.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n28\n\nBlei, D. M., Ng, A. Y., and Jordan, M. I. Latent dirichlet allocation. Journal of machine Learning research, 3(Jan): 993–1022, 2003.\n\nChen, X., Fan, H., Girshick, R., and He, K. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020d.\n\nBolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29:4349–4357, 2016.\nBowker, G. C. and Star, S. L. Sorting things out: Classiﬁcation and its consequences. MIT press, 2000.\nBrown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.\nBrowne, S. Dark Matters: Surveillance of Blackness. Duke University Press, 2015.\nBulent Sariyildiz, M., Perez, J., and Larlus, D. Learning visual representations with caption annotations. arXiv e-prints, pp. arXiv–2008, 2020.\nBuolamwini, J. and Gebru, T. Gender shades: Intersectional accuracy disparities in commercial gender classiﬁcation. In Conference on fairness, accountability and transparency, pp. 77–91, 2018.\nCarreira, J., Noland, E., Hillier, C., and Zisserman, A. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987, 2019.\nChen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. Generative pretraining from pixels. In International Conference on Machine Learning, pp. 1691–1703. PMLR, 2020a.\nChen, T., Xu, B., Zhang, C., and Guestrin, C. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174, 2016.\nChen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020b.\nChen, T., Kornblith, S., Swersky, K., Norouzi, M., and Hinton, G. Big self-supervised models are strong semisupervised learners. arXiv preprint arXiv:2006.10029, 2020c.\nChen, X. and Gupta, A. Webly supervised learning of convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1431– 1439, 2015.\n\nChen, Y.-C., Li, L., Yu, L., Kholy, A. E., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. Uniter: Learning universal imagetext representations. arXiv preprint arXiv:1909.11740, 2019.\nCheng, G., Han, J., and Lu, X. Remote sensing image scene classiﬁcation: Benchmark and state of the art. Proceedings of the IEEE, 105(10):1865–1883, 2017.\nChoi, D., Shallue, C. J., Nado, Z., Lee, J., Maddison, C. J., and Dahl, G. E. On empirical comparisons of optimizers for deep learning. arXiv preprint arXiv:1910.05446, 2019.\nCoates, A., Ng, A., and Lee, H. An analysis of singlelayer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artiﬁcial intelligence and statistics, pp. 215–223, 2011.\nCrawford, K. The trouble with bias. NIPS 2017 Keynote, 2017. URL https://www.youtube.com/ watch?v=fMym_BKWQzk.\nDai, A. M. and Le, Q. V. Semi-supervised sequence learning. In Advances in neural information processing systems, pp. 3079–3087, 2015.\nD’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al. Underspeciﬁcation presents challenges for credibility in modern machine learning. arXiv preprint arXiv:2011.03395, 2020.\nDeng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and FeiFei, L. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09, 2009.\nDeng, J., Berg, A. C., Satheesh, S., Su, H., Khosla, A., and Fei-Fei, L. Ilsvrc 2012, 2012. URL http://www. image-net.org/challenges/LSVRC/2012/.\nDesai, K. and Johnson, J. Virtex: Learning visual representations from textual annotations. arXiv preprint arXiv:2006.06666, 2020.\nDevlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.\nDhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n29\n\nDivvala, S. K., Farhadi, A., and Guestrin, C. Learning everything about anything: Webly-supervised visual concept learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3270– 3277, 2014.\nDodge, S. and Karam, L. A study and comparison of human and deep learning recognition performance under visual distortions. In 2017 26th international conference on computer communication and networks (ICCCN), pp. 1– 7. IEEE, 2017.\nDosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.\nElhoseiny, M., Saleh, B., and Elgammal, A. Write a classiﬁer: Zero-shot learning using purely textual descriptions. In Proceedings of the IEEE International Conference on Computer Vision, pp. 2584–2591, 2013.\nFaghri, F., Fleet, D. J., Kiros, J. R., and Fidler, S. Vse++: Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612, 2017.\nFergus, R., Fei-Fei, L., Perona, P., and Zisserman, A. Learning object categories from google’s image search. In Tenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, volume 2, pp. 1816–1823. IEEE, 2005.\nFrome, A., Corrado, G. S., Shlens, J., Bengio, S., Dean, J., Ranzato, M., and Mikolov, T. Devise: A deep visualsemantic embedding model. In Advances in neural information processing systems, pp. 2121–2129, 2013.\nGan, Z., Chen, Y.-C., Li, L., Zhu, C., Cheng, Y., and Liu, J. Large-scale adversarial training for vision-and-language representation learning. arXiv preprint arXiv:2006.06195, 2020.\nGao, T., Fisch, A., and Chen, D. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723, 2020.\nGarvie, C., May 2019. URL https://www. flawedfacedata.com/.\nGeiger, A., Lenz, P., and Urtasun, R. Are we ready for autonomous driving? the kitti vision benchmark suite. In Conference on Computer Vision and Pattern Recognition (CVPR), 2012.\nGeirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W. Imagenet-trained cnns are\n\nbiased towards texture; increasing shape bias improves accuracy and robustness. arXiv preprint arXiv:1811.12231, 2018.\nGeirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A. Shortcut learning in deep neural networks. arXiv preprint arXiv:2004.07780, 2020.\nGomez, L., Patel, Y., Rusin˜ol, M., Karatzas, D., and Jawahar, C. Self-supervised learning of visual features through embedding images into text topic spaces. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4230–4239, 2017.\nGoodfellow, I. J., Shlens, J., and Szegedy, C. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.\nGoodfellow, I. J., Erhan, D., Carrier, P. L., Courville, A., Mirza, M., Hamner, B., Cukierski, W., Tang, Y., Thaler, D., Lee, D.-H., et al. Challenges in representation learning: A report on three machine learning contests. Neural Networks, 64:59–63, 2015.\nGoogle. Google cloud api: Celebrity recognition. URL https://cloud.google.com/vision/docs/ celebrity-recognition.\nGriewank, A. and Walther, A. Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation. ACM Transactions on Mathematical Software (TOMS), 26(1):19–45, 2000.\nGrill, J.-B., Strub, F., Altche´, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.\nHa, D., Dai, A., and Le, Q. V. Hypernetworks. arXiv preprint arXiv:1609.09106, 2016.\nHancock, B., Bringmann, M., Varma, P., Liang, P., Wang, S., and Re´, C. Training classiﬁers with natural language explanations. In Proceedings of the conference. Association for Computational Linguistics. Meeting, volume 2018, pp. 1884. NIH Public Access, 2018.\nHancock, B., Bordes, A., Mazare, P.-E., and Weston, J. Learning from dialogue after deployment: Feed yourself, chatbot! arXiv preprint arXiv:1901.05415, 2019.\nHarris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., Ferna´ndez del\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n30\n\nR´ıo, J., Wiebe, M., Peterson, P., Ge´rard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E. Array programming with NumPy. Nature, 585:357–362, 2020. doi: 10.1038/ s41586-020-2649-2.\nHays, J. and Efros, A. A. Im2gps: estimating geographic information from a single image. In 2008 ieee conference on computer vision and pattern recognition, pp. 1–8. IEEE, 2008.\nHe, K., Zhang, X., Ren, S., and Sun, J. Delving deep into rectiﬁers: Surpassing human-level performance on imagenet classiﬁcation. In Proceedings of the IEEE international conference on computer vision, pp. 1026–1034, 2015.\nHe, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016a.\nHe, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016b.\nHe, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9729– 9738, 2020.\nHe, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., and Li, M. Bag of tricks for image classiﬁcation with convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 558– 567, 2019.\nHe, X. and Peng, Y. Fine-grained image classiﬁcation via combining vision and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5994–6002, 2017.\nHelber, P., Bischke, B., Dengel, A., and Borth, D. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classiﬁcation. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019.\nHenaff, O. Data-efﬁcient image recognition with contrastive predictive coding. In International Conference on Machine Learning, pp. 4182–4192. PMLR, 2020.\nHendrycks, D. and Dietterich, T. Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261, 2019.\n\nHendrycks, D. and Gimpel, K. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.\nHendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D. Natural adversarial examples. arXiv preprint arXiv:1907.07174, 2019.\nHendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. arXiv preprint arXiv:2006.16241, 2020a.\nHendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D. Pretrained transformers improve out-ofdistribution robustness. arXiv preprint arXiv:2004.06100, 2020b.\nHestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Ali, M., Yang, Y., and Zhou, Y. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409, 2017.\nHill, F., Lampinen, A., Schneider, R., Clark, S., Botvinick, M., McClelland, J. L., and Santoro, A. Environmental drivers of systematicity and generalization in a situated agent. In International Conference on Learning Representations, 2019.\nHodosh, M., Young, P., and Hockenmaier, J. Framing image description as a ranking task: Data, models and evaluation metrics. Journal of Artiﬁcial Intelligence Research, 47: 853–899, 2013.\nHongsuck Seo, P., Weyand, T., Sim, J., and Han, B. Cplanet: Enhancing image geolocalization by combinatorial partitioning of maps. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 536–551, 2018.\nHoward, J. and Ruder, S. Universal language model ﬁne-tuning for text classiﬁcation. arXiv preprint arXiv:1801.06146, 2018.\nIlyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A. Adversarial examples are not bugs, they are features. In Advances in Neural Information Processing Systems, pp. 125–136, 2019.\nIoffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.\nJaderberg, M., Simonyan, K., Vedaldi, A., and Zisserman, A. Deep structured output learning for unconstrained text recognition. arXiv preprint arXiv:1412.5903, 2014.\nJaderberg, M., Simonyan, K., Zisserman, A., et al. Spatial transformer networks. Advances in neural information processing systems, 28:2017–2025, 2015.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n31\n\nJohnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2901–2910, 2017.\n\nKrishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1):32–73, 2017.\n\nJoulin, A., Van Der Maaten, L., Jabri, A., and Vasilache, N. Learning visual features from large weakly supervised data. In European Conference on Computer Vision, pp. 67–84. Springer, 2016.\nKalfaoglu, M., Kalkan, S., and Alatan, A. A. Late temporal modeling in 3d cnn architectures with bert for action recognition. arXiv preprint arXiv:2008.01232, 2020.\nKaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.\nKarpathy, A., Joulin, A., and Fei-Fei, L. F. Deep fragment embeddings for bidirectional image sentence mapping. In Advances in neural information processing systems, pp. 1889–1897, 2014.\nKeyes, O. The misgendering machines: Trans/hci implications of automatic gender recognition. Proceedings of the ACM on Human-Computer Interaction, 2(CSCW):1–22, 2018.\nKiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., and Testuggine, D. The hateful memes challenge: Detecting hate speech in multimodal memes. arXiv preprint arXiv:2005.04790, 2020.\nKingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.\nKiros, R., Salakhutdinov, R., and Zemel, R. S. Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539, 2014.\nKiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S. Skip-thought vectors. Advances in neural information processing systems, 28: 3294–3302, 2015.\nKolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N. Large scale learning of general visual representations for transfer. arXiv preprint arXiv:1912.11370, 2019.\nKornblith, S., Shlens, J., and Le, Q. V. Do better imagenet models transfer better? In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2661–2671, 2019.\n\nKrizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classiﬁcation with deep convolutional neural networks. In Advances in neural information processing systems, pp. 1097–1105, 2012.\nKuhnle, A. and Copestake, A. Shapeworld-a new test methodology for multimodal language understanding. arXiv preprint arXiv:1704.04517, 2017.\nKa¨rkka¨inen, K. and Joo, J. Fairface: Face attribute dataset for balanced race, gender, and age, 2019.\nLake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. Building machines that learn and think like people, 2016.\nLampert, C. H., Nickisch, H., and Harmeling, S. Learning to detect unseen object classes by between-class attribute transfer. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 951–958. IEEE, 2009.\nLarochelle, H., Erhan, D., and Bengio, Y. Zero-data learning of new tasks. 2008.\nLe, Q. and Mikolov, T. Distributed representations of sentences and documents. In International conference on machine learning, pp. 1188–1196, 2014.\nLeCun, Y. The mnist database of handwritten digits. http://yann. lecun. com/exdb/mnist/.\nLee, D.-H. Pseudo-label: The simple and efﬁcient semisupervised learning method for deep neural networks.\nLei Ba, J., Swersky, K., Fidler, S., et al. Predicting deep zero-shot convolutional neural networks using textual descriptions. In Proceedings of the IEEE International Conference on Computer Vision, pp. 4247–4255, 2015.\nLi, A., Jabri, A., Joulin, A., and van der Maaten, L. Learning visual n-grams from web data. In Proceedings of the IEEE International Conference on Computer Vision, pp. 4183–4192, 2017.\nLi, G., Duan, N., Fang, Y., Gong, M., and Jiang, D. Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. 2020a.\nLi, J., Miller, A. H., Chopra, S., Ranzato, M., and Weston, J. Learning through dialogue interactions by asking questions. arXiv preprint arXiv:1612.04936, 2016.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n32\n\nLi, X., Yin, X., Li, C., Hu, X., Zhang, P., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., et al. Oscar: Objectsemantics aligned pre-training for vision-language tasks. arXiv preprint arXiv:2004.06165, 2020b.\n\nLiang, W., Zou, J., and Yu, Z. Alice: Active learning with contrastive natural language explanations. arXiv preprint arXiv:2009.10259, 2020.\n\nLin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dolla´r, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740–755. Springer, 2014.\n\nLinzen, T. How can we accelerate progress towards human-like linguistic generalization? arXiv preprint arXiv:2005.00955, 2020.\n\nLippe, P., Holla, N., Chandra, S., Rajamanickam, S., Antoniou, G., Shutova, E., and Yannakoudakis, H. A multimodal framework for the detection of hateful memes. arXiv preprint arXiv:2012.12871, 2020.\n\nLiu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N. Generating wikipedia by summarizing long sequences. arXiv preprint arXiv:1801.10198, 2018.\n\nLocatello, F., Bauer, S., Lucic, M., Ra¨tsch, G., Gelly, S., Scho¨lkopf, B., and Bachem, O. A sober look at the unsupervised learning of disentangled representations and their evaluation. arXiv preprint arXiv:2010.14766, 2020.\n\nLoshchilov, I. and Hutter, F. Sgdr: dient descent with warm restarts. arXiv:1608.03983, 2016.\n\nStochastic graarXiv preprint\n\nLoshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.\n\nLu, J., Batra, D., Parikh, D., and Lee, S. Vilbert: Pretraining task-agnostic visiolinguistic representations for visionand-language tasks. In Advances in Neural Information Processing Systems, pp. 13–23, 2019.\n\nLu, Z., Xiong, X., Li, Y., Stroud, J., and Ross, D. Leveraging weakly supervised data and pose representation for action recognition, 2020. URL https://www.youtube. com/watch?v=KOQFxbPPLOE&t=1390s.\n\nLucic, M., Kurach, K., Michalski, M., Gelly, S., and Bousquet, O. Are gans created equal? a large-scale study. Advances in neural information processing systems, 31: 700–709, 2018.\n\nMahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and van der Maaten, L. Exploring the limits of weakly supervised pretraining. In\n\nProceedings of the European Conference on Computer Vision (ECCV), pp. 181–196, 2018.\nMcCann, B., Bradbury, J., Xiong, C., and Socher, R. Learned in translation: Contextualized word vectors. In Advances in neural information processing systems, pp. 6294–6305, 2017.\nMcCann, B., Keskar, N. S., Xiong, C., and Socher, R. The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730, 2018.\nMicikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al. Mixed precision training. arXiv preprint arXiv:1710.03740, 2017.\nMiech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE international conference on computer vision, pp. 2630–2640, 2019.\nMiech, A., Alayrac, J.-B., Laptev, I., Sivic, J., and Zisserman, A. Rareact: A video dataset of unusual interactions. arXiv preprint arXiv:2008.01018, 2020a.\nMiech, A., Alayrac, J.-B., Smaira, L., Laptev, I., Sivic, J., and Zisserman, A. End-to-end learning of visual representations from uncurated instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9879–9889, 2020b.\nMikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, 26:3111–3119, 2013.\nMiller, J., Krauth, K., Recht, B., and Schmidt, L. The effect of natural distribution shift on question answering models. arXiv preprint arXiv:2004.14444, 2020.\nMishra, A., Alahari, K., and Jawahar, C. Scene text recognition using higher order language priors. 2012.\nMithun, N. C., Panda, R., Papalexakis, E. E., and RoyChowdhury, A. K. Webly supervised joint embedding for cross-modal image-text retrieval. In Proceedings of the 26th ACM international conference on Multimedia, pp. 1856–1864, 2018.\nMori, Y., Takahashi, H., and Oka, R. Image-to-word transformation based on dividing and vector quantizing images with words. Citeseer, 1999.\nMu, J., Liang, P., and Goodman, N. Shaping visual representations with language for few-shot classiﬁcation. arXiv preprint arXiv:1911.02683, 2019.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n33\n\nMuller-Budack, E., Pustu-Iren, K., and Ewerth, R. Geolocation estimation of photos using a hierarchical model and scene classiﬁcation. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 563–579, 2018.\nMurty, S., Koh, P. W., and Liang, P. Expbert: Representation engineering with natural language explanations. arXiv preprint arXiv:2005.01932, 2020.\nNarasimhan, K., Kulkarni, T., and Barzilay, R. Language understanding for text-based games using deep reinforcement learning. arXiv preprint arXiv:1506.08941, 2015.\nNetzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. Reading digits in natural images with unsupervised feature learning. 2011.\nNoble, S. U. Algorithms of oppression: How search engines reinforce racism. 2018.\nNosek, B. A., Banaji, M. R., and Greenwald, A. G. Harvesting implicit group attitudes and beliefs from a demonstration web site. Group Dynamics: Theory, Research, and Practice, 6(1):101, 2002.\nOh, S., Hoogs, A., Perera, A., Cuntoor, N., Chen, C.-C., Lee, J. T., Mukherjee, S., Aggarwal, J., Lee, H., Davis, L., et al. A large-scale benchmark dataset for event recognition in surveillance video. In CVPR 2011, pp. 3153–3160. IEEE, 2011.\nOliver, A., Odena, A., Raffel, C. A., Cubuk, E. D., and Goodfellow, I. Realistic evaluation of deep semi-supervised learning algorithms. Advances in neural information processing systems, 31:3235–3246, 2018.\nOord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.\nOrdonez, V., Kulkarni, G., and Berg, T. Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems, 24:1143–1151, 2011.\npandas development team, T. pandas-dev/pandas: Pandas, February 2020. URL https://doi.org/10. 5281/zenodo.3509134.\nParkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V. Cats and dogs. In IEEE Conference on Computer Vision and Pattern Recognition, 2012.\nPaszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L.,\n\nBai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, pp. 8024– 8035, 2019.\nPedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.\nPennington, J., Socher, R., and Manning, C. D. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp. 1532–1543, 2014.\nPeters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. Deep contextualized word representations. arXiv preprint arXiv:1802.05365, 2018.\nQi, D., Su, L., Song, J., Cui, E., Bharti, T., and Sacheti, A. Imagebert: Cross-modal pre-training with largescale weak-supervised image-text data. arXiv preprint arXiv:2001.07966, 2020.\nQuattoni, A., Collins, M., and Darrell, T. Learning visual representations using images with captions. In 2007 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1–8. IEEE, 2007.\nRadford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pretraining, 2018.\nRadford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.\nRaffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a uniﬁed text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.\nRaji, I. D., Gebru, T., Mitchell, M., Buolamwini, J., Lee, J., and Denton, E. Saving face: Investigating the ethical concerns of facial recognition auditing, 2020.\nRamanathan, V., Liang, P., and Fei-Fei, L. Video event understanding using natural language descriptions. In Proceedings of the IEEE International Conference on Computer Vision, pp. 905–912, 2013.\nRashtchian, C., Young, P., Hodosh, M., and Hockenmaier, J. Collecting image annotations using amazon’s mechanical turk. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk, pp. 139–147, 2010.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n34\n\nRecht, B., Roelofs, R., Schmidt, L., and Shankar, V. Do imagenet classiﬁers generalize to imagenet? arXiv preprint arXiv:1902.10811, 2019.\nSalimans, T. and Kingma, D. P. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. In Advances in neural information processing systems, pp. 901–909, 2016.\nScheuerman, M. K., Paul, J. M., and Brubaker, J. R. How computers see gender: An evaluation of gender classiﬁcation in commercial facial analysis services. Proceedings of the ACM on Human-Computer Interaction, 3(CSCW): 1–33, 2019.\nSchwemmer, C., Knight, C., Bello-Pardo, E. D., Oklobdzija, S., Schoonvelde, M., and Lockhart, J. W. Diagnosing gender bias in image recognition systems. Socius, 6: 2378023120967171, 2020.\nSennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015.\nShankar, V., Dave, A., Roelofs, R., Ramanan, D., Recht, B., and Schmidt, L. Do image classiﬁers generalize across time? arXiv preprint arXiv:1906.02168, 2019.\nSharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556– 2565, 2018.\nSingh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. Towards vqa models that can read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8317–8326, 2019.\nSocher, R. and Fei-Fei, L. Connecting modalities: Semisupervised segmentation and annotation of images using unaligned text corpora. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 966–973. IEEE, 2010.\nSocher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pp. 1631–1642, 2013.\nSocher, R., Karpathy, A., Le, Q. V., Manning, C. D., and Ng, A. Y. Grounded compositional semantics for ﬁnding and describing images with sentences. Transactions of the Association for Computational Linguistics, 2:207–218, 2014.\n\nSohn, K. Improved deep metric learning with multi-class n-pair loss objective. In Advances in neural information processing systems, pp. 1857–1865, 2016.\nSolaiman, I., Brundage, M., Clark, J., Askell, A., HerbertVoss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., McCain, M., Newhouse, A., Blazakis, J., McGufﬁe, K., and Wang, J. Release strategies and the social impacts of language models, 2019.\nSoomro, K., Zamir, A. R., and Shah, M. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.\nSpeer, R. ftfy. Zenodo, 2019. URL https://doi.org/ 10.5281/zenodo.2591652. Version 5.5.\nSrivastava, N. and Salakhutdinov, R. Multimodal learning with deep boltzmann machines. In NIPS, 2012.\nSrivastava, S., Labutov, I., and Mitchell, T. Joint concept learning and semantic parsing from natural language explanations. In Proceedings of the 2017 conference on empirical methods in natural language processing, pp. 1527–1536, 2017.\nStallkamp, J., Schlipsing, M., Salmen, J., and Igel, C. The German Trafﬁc Sign Recognition Benchmark: A multiclass classiﬁcation competition. In IEEE International Joint Conference on Neural Networks, pp. 1453–1460, 2011.\nStroud, J. C., Ross, D. A., Sun, C., Deng, J., Sukthankar, R., and Schmid, C. Learning video representations from textual web supervision. arXiv preprint arXiv:2007.14937, 2020.\nSzegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. Inception-v4, inception-resnet and the impact of residual connections on learning. arXiv preprint arXiv:1602.07261, 2016.\nTan, H. and Bansal, M. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019.\nTan, M. and Le, Q. V. Efﬁcientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946, 2019.\nTaori, R., Dave, A., Shankar, V., Carlini, N., Recht, B., and Schmidt, L. Measuring robustness to natural distribution shifts in image classiﬁcation. arXiv preprint arXiv:2007.00644, 2020.\nThomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n35\n\nTian, Y., Krishnan, D., and Isola, P. Contrastive multiview coding. arXiv preprint arXiv:1906.05849, 2019.\nTian, Y., Wang, Y., Krishnan, D., Tenenbaum, J. B., and Isola, P. Rethinking few-shot image classiﬁcation: a good embedding is all you need? arXiv preprint arXiv:2003.11539, 2020.\nTorralba, A., Fergus, R., and Freeman, W. T. 80 million tiny images: A large data set for nonparametric object and scene recognition. IEEE transactions on pattern analysis and machine intelligence, 30(11):1958–1970, 2008.\nTouvron, H., Vedaldi, A., Douze, M., and Je´gou, H. Fixing the train-test resolution discrepancy. In Advances in neural information processing systems, pp. 8252–8262, 2019.\nVaradarajan, J. and Odobez, J.-M. Topic models for scene analysis and abnormality detection. In 2009 IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops, pp. 1338–1345. IEEE, 2009.\nVaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.\nVeeling, B. S., Linmans, J., Winkens, J., Cohen, T., and Welling, M. Rotation equivariant CNNs for digital pathology. June 2018.\nVirtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, ˙I., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors. SciPy 1.0: Fundamental Algorithms for Scientiﬁc Computing in Python. Nature Methods, 17:261–272, 2020. doi: 10.1038/s41592-019-0686-2.\nVo, N., Jacobs, N., and Hays, J. Revisiting im2gps in the deep learning era. In Proceedings of the IEEE International Conference on Computer Vision, pp. 2621–2630, 2017.\nWang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.\nWang, H., Ge, S., Lipton, Z., and Xing, E. P. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems, pp. 10506–10518, 2019.\n\nWang, H., Lu, P., Zhang, H., Yang, M., Bai, X., Xu, Y., He, M., Wang, Y., and Liu, W. All you need is boundary: Toward arbitrary-shaped text spotting. In Proceedings of the AAAI Conference on Artiﬁcial Intelligence, volume 34, pp. 12160–12167, 2020.\n\nWang, J., Markert, K., and Everingham, M. Learning models for object recognition from natural language descriptions. In BMVC, volume 1, pp. 2, 2009.\n\nWeston, J., Bengio, S., and Usunier, N. Large scale image annotation: learning to rank with joint word-image embeddings. Machine learning, 81(1):21–35, 2010.\n\nWeston, J. E. Dialog-based language learning. In Advances in Neural Information Processing Systems, pp. 829–837, 2016.\n\nWeyand, T., Kostrikov, I., and Philbin, J. Planet-photo geolocation with convolutional neural networks. In European Conference on Computer Vision, pp. 37–55. Springer, 2016.\n\nWu, Y., Kirillov, A., Massa, F., Lo, W.-Y., and Girshick, R. Detectron2. https://github.com/ facebookresearch/detectron2, 2019.\n\nWu, Z., Xiong, Y., Yu, S., and Lin, D. Unsupervised feature learning via non-parametric instance-level discrimination. arXiv preprint arXiv:1805.01978, 2018.\n\nXie, Q., Luong, M.-T., Hovy, E., and Le, Q. V. Self-training with noisy student improves imagenet classiﬁcation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10687–10698, 2020.\n\ny Arcas, B. A., Mitchell, M., and Todorov,\n\nA.\n\nPhysiognomy’s new clothes.\n\n2017.\n\nURL\n\nhttps://medium.com/@blaisea/\n\nphysiognomys-new-clothes-f2d4b59fdd6a.\n\nYang, Z., Lu, Y., Wang, J., Yin, X., Florencio, D., Wang, L., Zhang, C., Zhang, L., and Luo, J. Tap: Text-aware pre-training for text-vqa and text-caption. arXiv preprint arXiv:2012.04638, 2020.\n\nYogatama, D., d’Autume, C. d. M., Connor, J., Kocisky, T., Chrzanowski, M., Kong, L., Lazaridou, A., Ling, W., Yu, L., Dyer, C., et al. Learning and evaluating general linguistic intelligence. arXiv preprint arXiv:1901.11373, 2019.\n\nYoung, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.\n\n\fLearning Transferable Visual Models From Natural Language Supervision\n\n36\n\nYu, F., Tang, J., Yin, W., Sun, Y., Tian, H., Wu, H., and Wang, H. Ernie-vil: Knowledge enhanced visionlanguage representations through scene graph. arXiv preprint arXiv:2006.16934, 2020.\nZeiler, M. D. and Fergus, R. Visualizing and understanding convolutional networks. In European conference on computer vision, pp. 818–833. Springer, 2014.\nZhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al. A large-scale study of representation learning with the visual task adaptation benchmark. arXiv preprint arXiv:1910.04867, 2019.\nZhang, R. Making convolutional networks shift-invariant again. arXiv preprint arXiv:1904.11486, 2019.\nZhang, Y., Jiang, H., Miura, Y., Manning, C. D., and Langlotz, C. P. Contrastive learning of medical visual representations from paired images and text. arXiv preprint arXiv:2010.00747, 2020.\nZuboff, S. Big other: surveillance capitalism and the prospects of an information civilization. Journal of Information Technology, 30(1):75–89, 2015."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified.",
        "Non-reference conclusion/acknowledgment paragraphs interleaved by column extraction were excluded."
      ]
    },
    {
      "Slug": "latent-diffusion",
      "Paper": "High-Resolution Image Synthesis with Latent Diffusion Models",
      "AtlasYear": 2021,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2112.10752",
      "PdfSha256": "46EDE043A8DC07CA1F0F445620523FE1AD8B2436BD83856A3835612A47E9F79E",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 681,
          "EndLine": 899,
          "PdfPages": [
            10,
            11,
            12,
            13
          ],
          "Text": "[1] Eirikur Agustsson and Radu Timofte. NTIRE 2017 challenge on single image super-resolution: Dataset and study. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2017, Honolulu, HI, USA, July 21-26, 2017, pages 1122–1131. IEEE Computer Society, 2017. 1\n[2] Martin Arjovsky, Soumith Chintala, and Le´on Bottou. Wasserstein gan, 2017. 3\n[3] Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high ﬁdelity natural image synthesis. In Int. Conf. Learn. Represent., 2019. 1, 2, 7, 8, 22, 28\n[4] Holger Caesar, Jasper R. R. Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 1822, 2018, pages 1209–1218. Computer Vision Foundation / IEEE Computer Society, 2018. 7, 20, 22\n[5] Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650, 2021. 9\n[6] Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In ICML, volume 119 of Proceedings of Machine Learning Research, pages 1691–1703. PMLR, 2020. 3\n[7] Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan. Wavegrad: Estimating gradients for waveform generation. In ICLR. OpenReview.net, 2021. 1\n[8] Lu Chi, Borui Jiang, and Yadong Mu. Fast fourier convolution. In NeurIPS, 2020. 8\n[9] Rewon Child. Very deep vaes generalize autoregressive models and can outperform them on images. CoRR, abs/2011.10650, 2020. 3\n[10] Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. CoRR, abs/1904.10509, 2019. 3\n[11] Bin Dai and David P. Wipf. Diagnosing and enhancing VAE models. In ICLR (Poster). OpenReview.net, 2019. 2, 3\n[12] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. IEEE Computer Society, 2009. 1, 5, 7, 22\n[13] Emily Denton. Ethical considerations of generative ai. AI for Content Creation Workshop, CVPR, 2021. 9\n[14] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. CoRR, abs/1810.04805, 2018. 7\n[15] Prafulla Dhariwal and Alex Nichol. Diffusion models beat gans on image synthesis. CoRR, abs/2105.05233, 2021. 1, 2, 3, 4, 6, 7, 8, 18, 22, 25, 26, 28\n\n[16] Sander Dieleman. Musings on typicality, 2020. 1, 3 [17] Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng,\nChang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: Mastering text-toimage generation via transformers. CoRR, abs/2105.13290, 2021. 6, 7 [18] Laurent Dinh, David Krueger, and Yoshua Bengio. Nice: Non-linear independent components estimation, 2015. 3 [19] Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using real NVP. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. 1, 3 [20] Alexey Dosovitskiy and Thomas Brox. Generating images with perceptual similarity metrics based on deep networks. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors, Adv. Neural Inform. Process. Syst., pages 658–666, 2016. 3 [21] Patrick Esser, Robin Rombach, Andreas Blattmann, and Bjo¨rn Ommer. Imagebart: Bidirectional context with multinomial diffusion for autoregressive image synthesis. CoRR, abs/2108.08827, 2021. 6, 7, 22 [22] Patrick Esser, Robin Rombach, and Bjo¨rn Ommer. A note on data biases in generative models. arXiv preprint arXiv:2012.02516, 2020. 9 [23] Patrick Esser, Robin Rombach, and Bjo¨rn Ommer. Taming transformers for high-resolution image synthesis. CoRR, abs/2012.09841, 2020. 2, 3, 4, 6, 7, 21, 22, 29, 34, 36 [24] Mary Anne Franks and Ari Ezra Waldman. Sex, lies, and videotape: Deep fakes and free speech delusions. Md. L. Rev., 78:892, 2018. 9 [25] Kevin Frans, Lisa B. Soros, and Olaf Witkowski. Clipdraw: Exploring text-to-drawing synthesis through languageimage encoders. ArXiv, abs/2106.14843, 2021. 3 [26] Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scenebased text-to-image generation with human priors. CoRR, abs/2203.13131, 2022. 6, 7, 16 [27] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial networks. CoRR, 2014. 1, 2 [28] Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville. Improved training of wasserstein gans, 2017. 3 [29] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Adv. Neural Inform. Process. Syst., pages 6626– 6637, 2017. 1, 5, 26 [30] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 1, 2, 3, 4, 6, 17 [31] Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high ﬁdelity image generation. CoRR, abs/2106.15282, 2021. 1, 3, 22\n\n10\n\n\f[32] Jonathan Ho and Tim Salimans. Classiﬁer-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 6, 7, 16, 22, 28, 37, 38\n[33] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. Image-to-image translation with conditional adversarial networks. In CVPR, pages 5967–5976. IEEE Computer Society, 2017. 3, 4\n[34] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. Image-to-image translation with conditional adversarial networks. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5967–5976, 2017. 4\n[35] Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier J. He´naff, Matthew M. Botvinick, Andrew Zisserman, Oriol Vinyals, and Joa˜o Carreira. Perceiver IO: A general architecture for structured inputs &outputs. CoRR, abs/2107.14795, 2021. 4\n[36] Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joa˜o Carreira. Perceiver: General perception with iterative attention. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 4651–4664. PMLR, 2021. 4, 5\n[37] Manuel Jahn, Robin Rombach, and Bjo¨rn Ommer. Highresolution complex scene synthesis with transformers. CoRR, abs/2105.06458, 2021. 20, 22, 27\n[38] Niharika Jain, Alberto Olmo, Sailik Sengupta, Lydia Manikonda, and Subbarao Kambhampati. Imperfect imaganation: Implications of gans exacerbating biases on facial data augmentation and snapchat selﬁe lenses. arXiv preprint arXiv:2001.09528, 2020. 9\n[39] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. CoRR, abs/1710.10196, 2017. 5, 6\n[40] Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In IEEE Conf. Comput. Vis. Pattern Recog., pages 4401– 4410, 2019. 1\n[41] T. Karras, S. Laine, and T. Aila. A style-based generator architecture for generative adversarial networks. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 5, 6\n[42] Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. CoRR, abs/1912.04958, 2019. 2, 6, 28\n[43] Dongjun Kim, Seungjae Shin, Kyungwoo Song, Wanmo Kang, and Il-Chul Moon. Score matching model for unbounded data score. CoRR, abs/2106.05527, 2021. 6\n[44] Durk P Kingma and Prafulla Dhariwal. Glow: Generative ﬂow with invertible 1x1 convolutions. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, 2018. 3\n\n[45] Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. CoRR, abs/2107.00630, 2021. 1, 3, 16\n[46] Diederik P. Kingma and Max Welling. Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR, 2014. 1, 3, 4, 29\n[47] Zhifeng Kong and Wei Ping. On fast sampling of diffusion probabilistic models. CoRR, abs/2106.00132, 2021. 3\n[48] Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In ICLR. OpenReview.net, 2021. 1\n[49] Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper R. R. Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari. The open images dataset V4: uniﬁed image classiﬁcation, object detection, and visual relationship detection at scale. CoRR, abs/1811.00982, 2018. 7, 20, 22\n[50] Tuomas Kynka¨a¨nniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Improved precision and recall metric for assessing generative models. CoRR, abs/1904.06991, 2019. 5, 26\n[51] Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dolla´r, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. CoRR, abs/1405.0312, 2014. 6, 7, 27\n[52] Yuqing Ma, Xianglong Liu, Shihao Bai, Le-Yi Wang, Aishan Liu, Dacheng Tao, and Edwin Hancock. Region-wise generative adversarial imageinpainting for large missing areas. ArXiv, abs/1909.12507, 2019. 9\n[53] Chenlin Meng, Yang Song, Jiaming Song, Jiajun Wu, JunYan Zhu, and Stefano Ermon. Sdedit: Image synthesis and editing with stochastic differential equations. CoRR, abs/2108.01073, 2021. 1\n[54] Lars M. Mescheder. On the convergence properties of GAN training. CoRR, abs/1801.04406, 2018. 3\n[55] Luke Metz, Ben Poole, David Pfau, and Jascha SohlDickstein. Unrolled generative adversarial networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. 3\n[56] Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. CoRR, abs/1411.1784, 2014. 4\n[57] Gautam Mittal, Jesse H. Engel, Curtis Hawthorne, and Ian Simon. Symbolic music generation with diffusion models. CoRR, abs/2103.16091, 2021. 1\n[58] Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Z. Qureshi, and Mehran Ebrahimi. Edgeconnect: Generative image inpainting with adversarial edge learning. ArXiv, abs/1901.00212, 2019. 9\n[59] Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. CoRR, abs/2112.10741, 2021. 6, 7, 16\n[60] Anton Obukhov, Maximilian Seitzer, Po-Wei Wu, Semen Zhydenko, Jonathan Kyl, and Elvis Yu-Jing Lin.\n\n11\n\n\fHigh-ﬁdelity performance metrics for generative models in pytorch, 2020. Version: 0.3.0, DOI: 10.5281/zenodo.4957738. 26, 27\n[61] Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and JunYan Zhu. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019. 4, 7\n[62] Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and JunYan Zhu. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 22\n[63] Gaurav Parmar, Dacheng Li, Kwonjoon Lee, and Zhuowen Tu. Dual contradistinctive generative autoencoder. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 823–832. Computer Vision Foundation / IEEE, 2021. 6\n[64] Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On buggy resizing libraries and surprising subtleties in ﬁd calculation. arXiv preprint arXiv:2104.11222, 2021. 26\n[65] David A. Patterson, Joseph Gonzalez, Quoc V. Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. CoRR, abs/2104.10350, 2021. 2\n[66] Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. CoRR, abs/2102.12092, 2021. 1, 2, 3, 4, 7, 21, 27\n[67] Ali Razavi, Aa¨ron van den Oord, and Oriol Vinyals. Generating diverse high-ﬁdelity images with VQ-VAE-2. In NeurIPS, pages 14837–14847, 2019. 1, 2, 3, 22\n[68] Scott E. Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. Generative adversarial text to image synthesis. In ICML, 2016. 4\n[69] Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In Proceedings of the 31st International Conference on International Conference on Machine Learning, ICML, 2014. 1, 4, 29\n[70] Robin Rombach, Patrick Esser, and Bjo¨rn Ommer. Network-to-network translation with conditional invertible neural networks. In NeurIPS, 2020. 3\n[71] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. Unet: Convolutional networks for biomedical image segmentation. In MICCAI (3), volume 9351 of Lecture Notes in Computer Science, pages 234–241. Springer, 2015. 2, 3, 4\n[72] Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J. Fleet, and Mohammad Norouzi. Image super-resolution via iterative reﬁnement. CoRR, abs/2104.07636, 2021. 1, 4, 8, 16, 22, 23, 27\n[73] Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P. Kingma. Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modiﬁcations. CoRR, abs/1701.05517, 2017. 1, 3\n[74] Dave Salvator. NVIDIA Developer Blog. https : / / developer . nvidia . com / blog / getting -\n\nimmediate- speedups- with- a100- tf32, 2020. 28\n[75] Robin San-Roman, Eliya Nachmani, and Lior Wolf. Noise estimation for generative diffusion models. CoRR, abs/2104.02600, 2021. 3\n[76] Axel Sauer, Kashyap Chitta, Jens Mu¨ller, and Andreas Geiger. Projected gans converge faster. CoRR, abs/2111.01007, 2021. 6\n[77] Edgar Scho¨nfeld, Bernt Schiele, and Anna Khoreva. A unet based discriminator for generative adversarial networks. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 8204–8213. Computer Vision Foundation / IEEE, 2020. 6\n[78] Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion400m: Open dataset of clip-ﬁltered 400 million image-text pairs, 2021. 6, 7\n[79] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Yoshua Bengio and Yann LeCun, editors, Int. Conf. Learn. Represent., 2015. 29, 43, 44, 45\n[80] Abhishek Sinha, Jiaming Song, Chenlin Meng, and Stefano Ermon. D2C: diffusion-denoising models for few-shot conditional generation. CoRR, abs/2106.06819, 2021. 3\n[81] Charlie Snell. Alien Dreams: An Emerging Art Scene. https : / / ml . berkeley . edu / blog / posts / clip-art/, 2021. [Online; accessed November-2021]. 2\n[82] Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. CoRR, abs/1503.03585, 2015. 1, 3, 4, 18\n[83] Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learning structured output representation using deep conditional generative models. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. 4\n[84] Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR. OpenReview.net, 2021. 3, 5, 6, 22\n[85] Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Scorebased generative modeling through stochastic differential equations. CoRR, abs/2011.13456, 2020. 1, 3, 4, 18\n[86] Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for modern deep learning research. In The Thirty-Fourth AAAI Conference on Artiﬁcial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artiﬁcial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artiﬁcial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 13693–13696. AAAI Press, 2020. 2\n\n12\n\n\f[87] [88]\n[89]\n[90] [91] [92] [93] [94] [95] [96] [97] [98] [99] [100]\n\nWei Sun and Tianfu Wu. Learning layout and style re-\n\nconﬁgurable gans for controllable image synthesis. CoRR,\n\nabs/2003.11571, 2020. 22, 27\n\nRoman Suvorov, Elizaveta Logacheva, Anton Mashikhin,\n\nAnastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov,\n\nNaejin Kong, Harshith Goka, Kiwoong Park, and Victor S.\n\nLempitsky. Resolution-robust large mask inpainting with\n\nfourier convolutions. ArXiv, abs/2109.07161, 2021. 8, 9,\n\n26, 32\n\nTristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. De-\n\nvon Hjelm, and Shikhar Sharma. Object-centric image gen-\n\neration from layouts. In Thirty-Fifth AAAI Conference on\n\nArtiﬁcial Intelligence, AAAI 2021, Thirty-Third Conference\n\non Innovative Applications of Artiﬁcial Intelligence, IAAI\n\n2021, The Eleventh Symposium on Educational Advances\n\nin Artiﬁcial Intelligence, EAAI 2021, Virtual Event, Febru-\n\nary 2-9, 2021, pages 2647–2655. AAAI Press, 2021. 20,\n\n22, 27\n\nPatrick Tinsley, Adam Czajka, and Patrick Flynn. This face\n\ndoes not exist... but it might be yours! identity leakage in\n\ngenerative models. In Proceedings of the IEEE/CVF Win-\n\nter Conference on Applications of Computer Vision, pages\n\n1320–1328, 2021. 9\n\nAntonio Torralba and Alexei A Efros. Unbiased look at\n\ndataset bias. In CVPR 2011, pages 1521–1528. IEEE, 2011.\n\n9\n\nArash Vahdat and Jan Kautz. NVAE: A deep hierarchical\n\nvariational autoencoder. In NeurIPS, 2020. 3\n\nArash Vahdat, Karsten Kreis, and Jan Kautz. Score-\n\nbased generative modeling in latent space. CoRR,\n\nabs/2106.05931, 2021. 2, 3, 5, 6\n\nAaron van den Oord, Nal Kalchbrenner, Lasse Espeholt,\n\nkoray kavukcuoglu, Oriol Vinyals, and Alex Graves. Con-\n\nditional image generation with pixelcnn decoders. In Ad-\n\nvances in Neural Information Processing Systems, 2016. 3\n\nAa¨ron van den Oord, Nal Kalchbrenner, and Koray\n\nKavukcuoglu. Pixel recurrent neural networks. CoRR,\n\nabs/1601.06759, 2016. 3\n\nAa¨ron van den Oord, Oriol Vinyals, and Koray\n\nKavukcuoglu. Neural discrete representation learning. In\n\nNIPS, pages 6306–6315, 2017. 2, 4, 29\n\nAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob\n\nUszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser,\n\nand Illia Polosukhin. Attention is all you need. In NIPS,\n\npages 5998–6008, 2017. 3, 4, 5, 7\n\nRivers Have Wings.\n\nTweet on Classiﬁer-free\n\nguidance for autoregressive models.\n\nhttps :\n\n/ / twitter . com / RiversHaveWings / status /\n\n1478093658716966912, 2022. 6\n\nThomas Wolf, Lysandre Debut, Victor Sanh, Julien Chau-\n\nmond, Clement Delangue, Anthony Moi, Pierric Cistac,\n\nTim Rault, Re´mi Louf, Morgan Funtowicz, and Jamie\n\nBrew. Huggingface’s transformers: State-of-the-art natural\n\nlanguage processing. CoRR, abs/1910.03771, 2019. 26\n\nZhisheng Xiao, Karsten Kreis, Jan Kautz, and Arash Vah-\n\ndat. VAEBM: A symbiosis between variational autoen-\n\ncoders and energy-based models. In 9th International Con-\n\n[101] [102] [103] [104] [105] [106] [107] [108] [109]\n\nference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. 6\nWilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using VQ-VAE and transformers. CoRR, abs/2104.10157, 2021. 3\nFisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. LSUN: construction of a large-scale image dataset using deep learning with humans in the loop. CoRR, abs/1506.03365, 2015. 5\nJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan, 2021. 3, 4\nJiahui Yu, Zhe L. Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S. Huang. Free-form image inpainting with gated convolution. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 4470–4479, 2019. 9\nK. Zhang, Jingyun Liang, Luc Van Gool, and Radu Timofte. Designing a practical degradation model for deep blind image super-resolution. ArXiv, abs/2103.14006, 2021. 23\nRichard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 3, 8, 19\nShengyu Zhao, Jianwei Cui, Yilun Sheng, Yue Dong, Xiao Liang, Eric I-Chao Chang, and Yan Xu. Large scale image completion via co-modulated generative adversarial networks. ArXiv, abs/2103.10428, 2021. 9 Bolei Zhou, A` gata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40:1452–1464, 2018. 8, 9, 26\nYufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. LAFITE: towards language-free training for text-to-image generation. CoRR, abs/2111.13792, 2021. 6, 7, 16"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "chain-of-thought",
      "Paper": "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models",
      "AtlasYear": 2022,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2201.11903",
      "PdfSha256": "7D9F878C23B460E4566AA4EC9201B1ABFB3B8FAEFB2B1356E411CB90FEF72A12",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 471,
          "EndLine": 566,
          "PdfPages": [
            10,
            11,
            12,
            13,
            14
          ],
          "Text": "Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al. 2022. Do as I can, not as I say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691.\nAida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards interpretable math word problem solving with operationbased formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, Minnesota. Association for Computational Linguistics.\nDaniel Andor, Luheng He, Kenton Lee, and Emily Pitler. 2019. Giving BERT a calculator: Finding operations and arguments with reading comprehension. EMNLP.\nJacob Andreas, Dan Klein, and Sergey Levine. 2018. Learning with latent language. NAACL.\nJacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732.\nBIG-bench collaboration. 2021. Beyond the imitation game: Measuring and extrapolating the capabilities of language models. In preparation.\nKaj Bostrom, Xinyu Zhao, Swarat Chaudhuri, and Greg Durrett. 2021. Flexible generation of natural language deductions. EMNLP.\nTom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. NeurIPS.\nJonathon Cai, Richard Shin, and Dawn Song. 2017. Making neural programming architectures generalize via recursion. ICLR.\nOana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-SNLI: Natural language inference with natural language explanations. NeurIPS.\nHoward Chen, Jacqueline He, Karthik Narasimhan, and Danqi Chen. 2022. Can rationalization improve robustness? NAACL.\nMark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.\nXinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou, Dawn Song, and Quoc V. Le. 2019. Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension. ICLR.\nTing-Rui Chiang and Yun-Nung Chen. 2019. Semantically-aligned equation generation for solving and reasoning math word problems. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2656–2668, Minneapolis, Minnesota. Association for Computational Linguistics.\n10\n\n\fPeter Clark, Oyvind Tafjord, and Kyle Richardson. 2020. Transformers as soft reasoners over language. IJCAI.\nKarl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training veriﬁers to solve math word problems. arXiv preprint arXiv:2110.14168.\nJacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL.\nHonghua Dong, Jiayuan Mao, Tian Lin, Chong Wang, Lihong Li, and Denny Zhou. 2019. Neural logic machines. ICLR.\nDheeru Dua, Sameer Singh, and Matt Gardner. 2020. Beneﬁts of intermediate annotations in reading comprehension. ACL.\nMor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? A question answering benchmark with implicit reasoning strategies. TACL.\nYuling Gu, Bhavana Dalvi Mishra, and Peter Clark. 2022. DREAM: Uncovering mental models behind language models. NAACL.\nBraden Hancock, Paroma Varma, Stephanie Wang, Martin Bringmann, Percy Liang, and Christopher Ré. 2018. Training classiﬁers with natural language explanations. ACL.\nPeter Hase and Mohit Bansal. 2022. When can models learn from explanations? a formal framework for understanding the roles of explanation data. ACL.\nDan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874.\nMohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. EMNLP.\nZhanming Jie, Jierui Li, and Wei Lu. 2022. Learning to reason deductively: Math word problem solving as complex relation extraction. arXiv preprint arXiv:2203.10316.\nJared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.\nRik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. NAACL.\nAndrew K. Lampinen, Ishita Dasgupta, Stephanie C.Y. Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L. McClelland, Jane X. Wang, and Felix Hill. 2022. Can language models learn from explanations in context? arXiv preprint arXiv:2204.02329.\nYihuai Lan, Lei Wang, Qiyuan Zhang, Yunshi Lan, Bing Tian Dai, Yan Wang, Dongxiang Zhang, and Ee-Peng Lim. 2021. MWPToolkit: An open-source framework for deep learning-based math word problem solvers. arXiv preprint arXiv:2109.00799.\nTeven Le Scao and Alexander Rush. 2021. How many data points is a prompt worth? NAACL.\nBrian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efﬁcient prompt tuning. EMNLP.\nIddo Lev, Bill MacCartney, Christopher Manning, and Roger Levy. 2004. Solving logic puzzles: From robust processing to precise semantics. Proceedings of the 2nd Workshop on Text Meaning and Interpretation.\nXiang Lisa Li and Percy Liang. 2021. Preﬁx-tuning: Optimizing continuous prompts for generation. ACL.\n11\n\n\fZhengzhong Liang, Steven Bethard, and Mihai Surdeanu. 2021. Explainable multi-hop verbal reasoning through internal monologue. NAACL.\nWang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. ACL.\nPengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.\nBodhisattwa Prasad Majumder, Oana-Maria Camburu, Thomas Lukasiewicz, and Julian McAuley. 2021. Rationale-inspired natural language explanations with commonsense. arXiv preprint arXiv:2106.13876.\nAna Marasovic´, Iz Beltagy, Doug Downey, and Matthew E Peters. 2022. Few-shot self-rationalization with natural language prompts. NAACL Findings.\nJoshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In ACL.\nShen Yun Miao, Chao Chun Liang, and Keh Yih Su. 2020. A diverse corpus for evaluating and developing English math word problem solvers. ACL.\nSewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.\nSharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020. WT5?! Training text-to-text models to explain their predictions. arXiv preprint arXiv:2004.14546.\nMaxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114.\nLong Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.\nArkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? NAACL.\nMatthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. NAACL.\nXinyu Pi, Qian Liu, Bei Chen, Morteza Ziyadi, Zeqi Lin, Yan Gao, Qiang Fu, Jian-Guang Lou, and Weizhu Chen. 2022. Reasoning like program executors. arXiv preprint arXiv:2201.11473.\nPiotr Pie˛kos, Mateusz Malinowski, and Henryk Michalewski. 2021. Measuring and improving BERT’s mathematical abilities by predicting the order of reasoning. ACL.\nJack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training Gopher. arXiv preprint arXiv:2112.11446.\nColin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a uniﬁed text-to-text transformer. Journal of Machine Learning Research, 21:1–67.\nDheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, and Yulia Tsvetkov. 2021. SelfExplain: A self-explaining architecture for neural text classiﬁers. EMNLP.\nNazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Explain yourself! Leveraging language models for commonsense reasoning. ACL.\n12\n\n\fQiu Ran, Yankai Lin, Peng Li, Jie Zhou, and Zhiyuan Liu. 2019. NumNet: Machine reading comprehension with numerical reasoning. EMNLP.\nHannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021. Measuring attribution in natural language generation models. arXiv preprint arXiv:2112.12870.\nGabriel Recchia. 2021. Teaching autoregressive language models complex tasks by demonstration. arXiv preprint arXiv:2109.02102.\nEmily Reif, Daphne Ippolito, Ann Yuan, Andy Coenen, Chris Callison-Burch, and Jason Wei. 2022. A recipe for arbitrary text style transfer with large language models. ACL.\nLaria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems.\nSubhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. EMNLP.\nSubhro Roy, Tim Vieira, and Dan Roth. 2015. Reasoning about Quantities in Natural Language. TACL.\nMohammed Saeed, Naser Ahmadi, Preslav Nakov, and Paolo Papotti. 2021. RuleBERT: Teaching soft rules to pre-trained language models. EMNLP.\nVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chafﬁn, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2022. Multitask prompted training enables zero-shot task generalization. ICLR.\nJianhao Shen, Yichun Yin, Lin Li, Lifeng Shang, Xin Jiang, Ming Zhang, and Qun Liu. 2021. Generate & rank: A multi-task framework for math word problems. In Findings of the Association for Computational Linguistics: EMNLP 2021.\nAlon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. NAACL.\nAlon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, and Jonathan Berant. 2020. Leap-ofthought: Teaching pre-trained models to systematically reason over implicit knowledge. NeurIPS.\nAlon Talmor, Ori Yoran, Ronan Le Bras, Chandra Bhagavatula, Yoav Goldberg, Yejin Choi, and Jonathan Berant. 2021. CommonsenseQA 2.0: Exposing the limits of ai through gamiﬁcation. NeurIPS Track on Datasets and Benchmarks.\nYi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131.\nRomal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. LaMDA: Language models for dialog applications. arXiv preprint arXiv:2201.08239.\nXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022a. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.\nYizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022b. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv preprint arXiv:2204.07705.\nJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a. Finetuned language models are zero-shot learners. ICLR.\n13\n\n\fJason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022b. Emergent abilities of large language models. Transactions on Machine Learning Research.\nSarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi. 2022. Reframing human-AI collaboration for generating free-text explanations. NAACL.\nSarah Wiegreffe and Ana Marasovic´. 2021. Teach me to explain: A review of datasets for explainable NLP. NeurIPS.\nSarah Wiegreffe, Ana Marasovic´, and Noah A. Smith. 2021. Measuring association between labels and free-text rationales. EMNLP.\nTongshuang Wu, Ellen Jiang, Aaron Donsbach, Jeff Gray, Alejandra Molina, Michael Terry, and Carrie J Cai. 2022a. PromptChainer: Chaining large language model prompts through visual programming. CHI Extended Abstracts.\nTongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022b. AI chains: Transparent and controllable human-AI interaction by chaining large language model prompts. CHI.\nYujun Yan, Kevin Swersky, Danai Koutra, Parthasarathy Ranganathan, and Milad Hashemi. 2020. Neural execution engines: Learning to execute subroutines. NeurIPS.\nHuihan Yao, Ying Chen, Qinyuan Ye, Xisen Jin, and Xiang Ren. 2021. Reﬁning language models with compositional explanations. NeurIPS.\nXi Ye and Greg Durrett. 2022. The unreliability of explanations in few-shot in-context learning. arXiv preprint arXiv:2205.03401.\nYordan Yordanov, Vid Kocijan, Thomas Lukasiewicz, and Oana-Maria Camburu. 2021. Few-shot out-of-domain transfer learning of natural language explanations. arXiv preprint arXiv:2112.06204.\nOmar Zaidan, Jason Eisner, and Christine Piatko. 2007. Using “annotator rationales” to improve machine learning for text categorization. NAACL.\nWojciech Zaremba and Ilya Sutskever. 2014. Learning to execute. arXiv preprint arXiv:1410.4615. Eric Zelikman, Yuhuai Wu, and Noah D. Goodman. 2022. STaR: Bootstrapping reasoning with\nreasoning. arXiv preprint arXiv:2203.14465. Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use:\nImproving few-shot performance of language models. ICML. Wangchunshu Zhou, Jinyi Hu, Hanlin Zhang, Xiaodan Liang, Maosong Sun, Chenyan Xiong, and\nJian Tang. 2020. Towards interpretable natural language understanding with explanations as latent variables. NeurIPS.\n14"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "instructgpt",
      "Paper": "Training language models to follow instructions with human feedback",
      "AtlasYear": 2022,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2203.02155",
      "PdfSha256": "C1984BB50A5B90FDDB895FDC3A0F72E5BC977148C9F63EF6040CBE7A3E1F0D98",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 506,
          "EndLine": 606,
          "PdfPages": [
            21,
            22,
            23,
            24,
            25
          ],
          "Text": "Abramson, J., Ahuja, A., Barr, I., Brussee, A., Carnevale, F., Cassin, M., Chhaparia, R., Clark, S., Damoc, B., Dudzik, A., et al. (2020). Imitating interactive intelligence. arXiv preprint arXiv:2012.05672.\nAchiam, J., Held, D., Tamar, A., and Abbeel, P. (2017). Constrained policy optimization. In International Conference on Machine Learning, pages 22–31. PMLR.\nAnthony, T., Tian, Z., and Barber, D. (2017). Thinking fast and slow with deep learning and tree search. arXiv preprint arXiv:1705.08439.\nAribandi, V., Tay, Y., Schuster, T., Rao, J., Zheng, H. S., Mehta, S. V., Zhuang, H., Tran, V. Q., Bahri, D., Ni, J., et al. (2021). Ext5: Towards extreme multi-task scaling for transfer learning. arXiv preprint arXiv:2111.10952.\nAskell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. (2021). A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861.\nBahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A., and Bengio, Y. (2016). An actor-critic algorithm for sequence prediction. arXiv preprint arXiv:1607.07086.\nBahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E. (2018). Learning to understand goal speciﬁcations by modelling reward. arXiv preprint arXiv:1806.01946.\nBender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.\nBlodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H. (2020). Language (technology) is power: A critical survey of\" bias\" in nlp. arXiv preprint arXiv:2005.14050.\n21\n\n\fBöhm, F., Gao, Y., Meyer, C. M., Shapira, O., Dagan, I., and Gurevych, I. (2019). Better rewards yield better summaries: Learning to summarise without references. arXiv preprint arXiv:1909.01214.\nBojar, O., Chatterjee, R., Federmann, C., Haddow, B., Huck, M., Hokamp, C., Koehn, P., Logacheva, V., Monz, C., Negri, M., Post, M., Scarton, C., Specia, L., and Turchi, M. (2015). Findings of the 2015 workshop on statistical machine translation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 1–46, Lisbon, Portugal. Association for Computational Linguistics.\nBommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.\nBostrom, N. (2014). Superintelligence. Dunod.\nBrown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. arXiv preprint arXiv:2005.14165.\nBuchanan, B., Lohn, A., Musser, M., and Sedova, K. (2021). Truth, lies, and automation. Technical report, Center for the Study of Emerging Technology.\nCaliskan, A., Bryson, J. J., and Narayanan, A. (2017). Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.\nCarlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al. (2021). Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650.\nChen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.\nCho, W. S., Zhang, P., Zhang, Y., Li, X., Galley, M., Brockett, C., Wang, M., and Gao, J. (2018). Towards coherent and cohesive long-form text generation. arXiv preprint arXiv:1811.00511.\nChoi, E., He, H., Iyyer, M., Yatskar, M., Yih, W.-t., Choi, Y., Liang, P., and Zettlemoyer, L. (2018). Quac: Question answering in context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2174–2184.\nChristiano, P., Cotra, A., and Xu, M. (2021). Eliciting latent knowledge: How to tell if your eyes deceive you. https://www.alignmentforum.org/posts/qHCDysDnvhteW7kRd/arc-s-ﬁrst-technicalreport-eliciting-latent-knowledge.\nChristiano, P., Shlegeris, B., and Amodei, D. (2018). Supervising strong learners by amplifying weak experts. arXiv preprint arXiv:1810.08575.\nChristiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, pages 4299–4307.\nDathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. (2019). Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164.\nDhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K.-W., and Gupta, R. (2021). Bold: Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 862–872.\nDinan, E., Fan, A., Williams, A., Urbanek, J., Kiela, D., and Weston, J. (2019a). Queens are powerful too: Mitigating gender bias in dialogue generation. arXiv preprint arXiv:1911.03842.\nDinan, E., Humeau, S., Chintagunta, B., and Weston, J. (2019b). Build it break it ﬁx it for dialogue safety: Robustness from adversarial human attack. arXiv preprint arXiv:1908.06083.\nDua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M. (2019). Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. arXiv preprint arXiv:1903.00161.\nFedus, W., Zoph, B., and Shazeer, N. (2021). Switch transformers: Scaling to trillion parameter models with simple and efﬁcient sparsity. arXiv preprint arXiv:2101.03961.\n22\n\n\fGabriel, I. (2020). Artiﬁcial intelligence, values, and alignment. Minds and machines, 30(3):411–437.\nGehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. (2020). Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.\nHancock, B., Bordes, A., Mazare, P.-E., and Weston, J. (2019). Learning from dialogue after deployment: Feed yourself, chatbot! arXiv preprint arXiv:1901.05415.\nHenderson, P., Sinha, K., Angelard-Gontier, N., Ke, N. R., Fried, G., Lowe, R., and Pineau, J. (2018). Ethical challenges in data-driven dialogue systems. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, pages 123–129.\nHuang, P.-S., Zhang, H., Jiang, R., Stanforth, R., Welbl, J., Rae, J., Maini, V., Yogatama, D., and Kohli, P. (2019). Reducing sentiment bias in language models via counterfactual evaluation. arXiv preprint arXiv:1911.03064.\nIbarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D. (2018). Reward learning from human preferences and demonstrations in atari. In Advances in neural information processing systems, pages 8011–8023.\nIrving, G., Christiano, P., and Amodei, D. (2018). AI safety via debate. arXiv preprint arXiv:1805.00899.\nJaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019). Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456.\nKenton, Z., Everitt, T., Weidinger, L., Gabriel, I., Mikulik, V., and Irving, G. (2021). Alignment of language agents. arXiv preprint arXiv:2103.14659.\nKeskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. (2019). Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.\nKhashabi, D., Min, S., Khot, T., Sabharwal, A., Tafjord, O., Clark, P., and Hajishirzi, H. (2020). Uniﬁedqa: Crossing format boundaries with a single qa system. arXiv preprint arXiv:2005.00700.\nKirk, H., Jun, Y., Iqbal, H., Benussi, E., Volpin, F., Dreyer, F. A., Shtedritski, A., and Asano, Y. M. (2021). How true is gpt-2? an empirical analysis of intersectional occupational biases. arXiv preprint arXiv:2102.04130.\nKrause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. (2020). Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367.\nKreutzer, J., Khadivi, S., Matusov, E., and Riezler, S. (2018). Can neural machine translation be improved with user feedback? arXiv preprint arXiv:1804.05958.\nLawrence, C. and Riezler, S. (2018). Improving a neural semantic parser by counterfactual learning from human bandit feedback. arXiv preprint arXiv:1805.01252.\nLeike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. (2018). Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871.\nLeike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S. (2017). AI safety gridworlds. arXiv preprint arXiv:1711.09883.\nLiang, P. P., Wu, C., Morency, L.-P., and Salakhutdinov, R. (2021). Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.\nLin, S., Hilton, J., and Evans, O. (2021). Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.\nLiu, H., Dacon, J., Fan, W., Liu, H., Liu, Z., and Tang, J. (2019). Does gender matter? towards fairness in dialogue systems. arXiv preprint arXiv:1910.10486.\nMadaan, A., Tandon, N., Clark, P., and Yang, Y. (2022). Memory-assisted prompt editing to improve gpt-3 after deployment. arXiv preprint arXiv:2201.06009.\nManela, D. d. V., Errington, D., Fisher, T., van Breugel, B., and Minervini, P. (2021). Stereotype and skew: Quantifying gender bias in pre-trained and ﬁne-tuned language models. arXiv preprint arXiv:2101.09688.\nMishra, S., Khashabi, D., Baral, C., and Hajishirzi, H. (2021). Cross-task generalization via natural language crowdsourcing instructions. arXiv preprint arXiv:2104.08773.\n23\n\n\fNadeem, M., Bethke, A., and Reddy, S. (2020). Stereoset: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456.\nNahian, M. S. A., Frazier, S., Harrison, B., and Riedl, M. (2021). Training value-aligned reinforcement learning agents using a normative prior. arXiv preprint arXiv:2104.09469.\nNakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. (2021). Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.\nNallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al. (2016). Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023.\nNangia, N., Vania, C., Bhalerao, R., and Bowman, S. R. (2020). CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online. Association for Computational Linguistics.\nNgo, H., Raterink, C., Araújo, J. G., Zhang, I., Chen, C., Morisot, A., and Frosst, N. (2021). Mitigating harm in language models with conditional-likelihood ﬁltration. arXiv preprint arXiv:2108.07790.\nPerez, E., Karamcheti, S., Fergus, R., Weston, J., Kiela, D., and Cho, K. (2019). Finding generalizable evidence by learning to convince q&a models. arXiv preprint arXiv:1909.05863.\nQian, Y., Muaz, U., Zhang, B., and Hyun, J. W. (2019). Reducing gender bias in word-level language models with a gender-equalizing loss function. arXiv preprint arXiv:1905.12801.\nRadford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9.\nRae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. (2021). Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.\nRajpurkar, P., Jia, R., and Liang, P. (2018). Know what you don’t know: Unanswerable questions for squad. arXiv preprint arXiv:1806.03822.\nRudinger, R., Naradowsky, J., Leonard, B., and Van Durme, B. (2018). Gender bias in coreference resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, New Orleans, Louisiana. Association for Computational Linguistics.\nSanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chafﬁn, A., Stiegler, A., Scao, T. L., Raja, A., et al. (2021). Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.\nSchick, T., Udupa, S., and Schütze, H. (2021). Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. arXiv preprint arXiv:2103.00453.\nSchulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016). High-dimensional continuous control using generalized advantage estimation. In Proceedings of the International Conference on Learning Representations (ICLR).\nSchulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.\nSheng, E., Chang, K.-W., Natarajan, P., and Peng, N. (2019). The woman worked as a babysitter: On biases in language generation. arXiv preprint arXiv:1909.01326.\nSilver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2017). Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:1712.01815.\nSoares, N., Fallenstein, B., Armstrong, S., and Yudkowsky, E. (2015). Corrigibility. In Workshops at the Twenty-Ninth AAAI Conference on Artiﬁcial Intelligence.\nSocher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. (2013). Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.\n24\n\n\fSolaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., et al. (2019). Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203.\nSolaiman, I. and Dennison, C. (2021). Process for adapting language models to society (palms) with values-targeted datasets. arXiv preprint arXiv:2106.10328.\nStiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. (2020). Learning to summarize from human feedback. arXiv preprint arXiv:2009.01325.\nTamkin, A., Brundage, M., Clark, J., and Ganguli, D. (2021). Understanding the capabilities, limitations, and societal impact of large language models. arXiv preprint arXiv:2102.02503.\nThoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. (2022). Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.\nVig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. M. (2020). Investigating gender bias in language models using causal mediation analysis. In NeurIPS.\nVölske, M., Potthast, M., Syed, S., and Stein, B. (2017). Tl; dr: Mining reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization, pages 59–63.\nWang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2019). Superglue: A stickier benchmark for general-purpose language understanding systems. arXiv preprint arXiv:1905.00537.\nWei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. (2021). Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.\nWeidinger, L., Mellor, J., Rauh, M., Grifﬁn, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., et al. (2021). Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.\nWelbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S. (2021). Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445.\nWu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P. (2021). Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862.\nXu, A., Pathak, E., Wallace, E., Gururangan, S., Sap, M., and Klein, D. (2021). Detoxifying language models risks marginalizing minority voices. arXiv preprint arXiv:2104.06390.\nXu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. (2020). Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079.\nYi, S., Goel, R., Khatri, C., Cervone, A., Chung, T., Hedayatnia, B., Venkatesh, A., Gabriel, R., and Hakkani-Tur, D. (2019). Towards coherent and engaging spoken dialog response generation using automatic conversation evaluators. arXiv preprint arXiv:1904.13015.\nZellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019). Hellaswag: Can a machine really ﬁnish your sentence? In Association for Computational Linguistics, pages 4791–4800.\nZhao, M., Anderson, P., Jain, V., Wang, S., Ku, A., Baldridge, J., and Ie, E. (2021). On the evaluation of vision-and-language navigation instructions. arXiv preprint arXiv:2101.10504.\nZhou, W. and Xu, K. (2020). Learning to compare for better training and evaluation of open domain natural language generation models. arXiv preprint arXiv:2002.05058.\nZiegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019). Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593.\n25"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "zero-shot-reasoning",
      "Paper": "Large Language Models are Zero-Shot Reasoners",
      "AtlasYear": 2022,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2205.11916",
      "PdfSha256": "43A3D73C77C7F3E115BB85522A4D94AAA54CDB55A0D9E17EB6AF3117F39111C7",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 419,
          "EndLine": 477,
          "PdfPages": [
            10,
            11,
            12,
            13,
            14
          ],
          "Text": "Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian\n10\n\n\fIbarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, and Mengyuan Yan. Do as i can, not as i say: Grounding language in robotic affordances, 2022. URL https://arxiv.org/abs/ 2204.01691.\nSid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorﬂow, March 2021. URL https://doi. org/10.5281/zenodo.5297715.\nTom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in NeurIPS, volume 33, pages 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/ 1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.\nFrançois Chollet. On the measure of intelligence. arXiv preprint arXiv:1911.01547, 2019. URL https://arxiv.org/abs/1911.01547.\nAakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways, 2022. URL https://arxiv.org/abs/2204.02311.\nKarl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training veriﬁers to solve math word problems, 2021. URL https://arxiv.org/ abs/2110.14168.\nJacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL, pages 4171–4186, 2019. URL https://aclanthology.org/N19-1423.\nLeo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv: Arxiv-2101.00027, 2020.\nTianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. In Proceedings of ACL-IJCNLP, pages 3816–3830, 2021. URL https://aclanthology. org/2021.acl-long.295.\nMor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. TACL, 9:346–361, 2021. URL https://aclanthology.org/2021.tacl-1.21/.\nMohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. Learning to solve arithmetic word problems with verb categorization. In EMNLP, volume 523533. Citeseer, 2014. URL https://aclanthology.org/D14-1058/.\n11\n\n\fWendy Johnson and Thomas J Bouchard Jr. The structure of human intelligence: It is verbal, perceptual, and image rotation (vpr), not ﬂuid and crystallized. Intelligence, 33(4):393–416, 2005.\nRik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. Parsing algebraic word problems into equations. TACL, 3:585–597, 2015. URL https: //aclanthology.org/Q15-1042.\nRik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. MAWPS: A math word problem repository. In Proceedings of NAACL, pages 1152–1157, 2016. URL https://aclanthology.org/N16-1136.\nWang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of ACL, pages 158–167, 2017. URL https://aclanthology.org/P17-1015.\nJiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804, 2021a. URL https://arxiv.org/abs/2101.06804.\nPengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586, 2021b. URL https://arxiv.org/abs/2107. 13586.\nYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to ﬁnd them: Overcoming few-shot prompt order sensitivity. In Proceedings of ACL, pages 8086–8098, 2022. URL https://aclanthology.org/2022.acl-long.556.\nKevin S McGrew. The cattell-horn-carroll theory of cognitive abilities: Past, present, and future. 2005.\nStephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. arXiv preprint arXiv: Arxiv-1609.07843, 2016. URL https://arxiv.org/abs/1609. 07843.\nSewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837, 2022. URL https://arxiv.org/pdf/2202.12837.pdf.\nMaxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. Show your work: Scratchpads for intermediate computation with language models. In Deep Learning for Code Workshop, 2022. URL https://openreview.net/forum? id=HBlx2idbkbq.\nLong Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022. URL https://arxiv.org/abs/2203.02155.\nAdam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in NeurIPS, 32:8026–8037, 2019. URL https://papers.nips.cc/paper/2019/hash/ bdbca288fee7f92f2bfa9f7012727740-Abstract.html.\nArkil Patel, Satwik Bhattamishra, and Navin Goyal. Are NLP models really able to solve simple math word problems? In Proceedings of NAACL, pages 2080–2094, 2021. URL https:// aclanthology.org/2021.naacl-main.168.\nAlec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, page 9, 2019. URL http://www. persagen.com/files/misc/radford2019language.pdf.\n12\n\n\fJack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training gopher, 2021. URL https://arxiv.org/abs/2112.11446.\nColin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a uniﬁed text-to-text transformer. JMLR, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.\nNazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. Explain yourself! leveraging language models for commonsense reasoning. In Proceedings of ACL, pages 4932–4942, 2019. URL https://aclanthology.org/P19-1487.\nLaria Reynolds and Kyle McDonell. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–7, 2021. URL https://arxiv.org/pdf/2102.07350.pdf.\nSubhro Roy and Dan Roth. Solving general arithmetic word problems. In Proceedings of EMNLP, pages 1743–1752, 2015. URL https://aclanthology.org/D15-1202.\nVictor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chafﬁn, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. Multitask prompted training enables zero-shot task generalization. In ICLR, 2022. URL https://openreview.net/forum?id=9Vrb9D0WI4.\nTimo Schick and Hinrich Schütze. It’s not just size that matters: Small language models are also fewshot learners. In Proceedings of NAACL, pages 2339–2352, 2021. URL https://aclanthology. org/2021.naacl-main.185.\nTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of EMNLP, pages 4222–4235, 2020. URL https://aclanthology.org/2020. emnlp-main.346.\nVered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Unsupervised commonsense question answering with self-talk. In Proceedings of EMNLP, pages 4615–4629, 2020. URL https://aclanthology.org/2020.emnlp-main.373.\nShaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model, 2022. URL https: //arxiv.org/abs/2201.11990.\n13\n\n\fAarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022. URL https://arxiv.org/abs/2206.04615.\nKeith E Stanovich and Richard F West. Individual differences in reasoning: Implications for the rationality debate? Behavioral and brain sciences, 23(5):645–665, 2000.\nAlon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of NAACL-HLT, pages 4149–4158, 2019. URL https://aclanthology.org/N19-1421/.\nRomal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. Lamda: Language models for dialog applications, 2022. URL https: //arxiv.org/abs/2201.08239.\nAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in NeurIPS, 2017. URL https://proceedings.neurips.cc/paper/2017/file/ 3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.\nBen Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.\nXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022. URL https://arxiv.org/abs/2203.11171.\nAlbert Webson and Ellie Pavlick. Do prompt-based models really understand the meaning of their prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344. Association for Computational Linguistics, July 2022. URL https://aclanthology.org/2022.naacl-main. 167.\nJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models, 2022. URL https: //arxiv.org/abs/2201.11903.\nThomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP, 2020. URL https://aclanthology.org/ 2020.emnlp-demos.6.\nEric Zelikman, Yuhuai Wu, and Noah D. Goodman. Star: Bootstrapping reasoning with reasoning, 2022. URL https://arxiv.org/abs/2203.14465.\nSusan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022. URL https://arxiv.org/abs/2205.01068.\n14"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "constitutional-ai",
      "Paper": "Constitutional AI: Harmlessness from AI Feedback",
      "AtlasYear": 2022,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2212.08073",
      "PdfSha256": "9A456A07AD346E3372F9867D346F69F5B0F68B4C65F060ACA0B8A13FA9D98E83",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 588,
          "EndLine": 615,
          "PdfPages": [
            17,
            18
          ],
          "Text": "[Askell et al., 2021] Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatﬁeld-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., and Kaplan, J. (2021). A general language assistant as a laboratory for alignment.\n[Bai et al., 2022] Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatﬁeld-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J. (2022). Training a helpful and harmless assistant with reinforcement learning from human feedback.\n[Bowman et al., 2022] Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukosuite, K., Askell, A., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Olah, C., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Kernion, J., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lovitt, L., Elhage, N., Schiefer, N., Joseph, N., Mercado, N., DasSarma, N., Larson, R., McCandlish, S., Kundu, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatﬁeld-Dodds, Z., Mann, B., and Kaplan, J. (2022). Measuring progress on scalable oversight for large language models.\n[Christiano et al., 2017] Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences.\n[Christiano et al., 2018] Christiano, P., Shlegeris, B., and Amodei, D. (2018). Supervising strong learners by amplifying weak experts.\n[Ganguli et al., 2022] Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Dodds, Z. H., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned.\n[Gao et al., 2022] Gao, L., Schulman, J., and Hilton, J. (2022). Scaling laws for reward model overoptimization.\n[Glaese et al., 2022] Glaese, A., McAleese, N., Tre˛bacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., MokrÃ¡, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G. (2022). Improving alignment of dialogue agents via targeted human judgements.\n[Huang et al., 2022] Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J. (2022). Large language models can self-improve.\n[Irving et al., 2018] Irving, G., Christiano, P., and Amodei, D. (2018). Ai safety via debate.\n[Kadavath et al., 2022] Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Dodds, Z. H., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J. (2022). Language models (mostly) know what they know.\n[Kojima et al., 2022] Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. (2022). Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.\n17\n\n\f[Nye et al., 2021] Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A. (2021). Show your work: Scratchpads for intermediate computation with language models.\n[Ouyang et al., 2022] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.\n[Perez et al., 2022] Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. (2022). Red teaming language models with language models.\n[Saunders et al., 2022] Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. (2022). Self-critiquing models for assisting human evaluators.\n[Scheurer et al., ] Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K., and Perez, E. Training language models with language feedback.\n[Shi et al., 2022] Shi, W., Dinan, E., Shuster, K., Weston, J., and Xu, J. (2022). When life gives you lemons, make cherryade: Converting feedback from bad responses into good labels.\n[Silver et al., 2017] Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2017). Mastering chess and shogi by self-play with a general reinforcement learning algorithm.\n[Solaiman and Dennison, 2021] Solaiman, I. and Dennison, C. (2021). Process for adapting language models to society (PALMS) with values-targeted datasets. CoRR, abs/2106.10328.\n[Srivastava et al., 2022] Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2022). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.\n[Stiennon et al., 2020] Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. (2020). Learning to summarize from human feedback.\n[Thoppilan et al., 2022] Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C., Krivokon, I., Rusch, W., Pickett, M., Meier-Hellstern, K. S., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Soraker, J., Zevenbergen, B., Prabhakaran, V., Diaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., Chi, E., and Le, Q. (2022). Lamda: Language models for dialog applications. CoRR, abs/2201.08239.\n[Wei et al., 2022] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. (2022). Chain of thought prompting elicits reasoning in large language models.\n[Xu et al., 2020] Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. (2020). Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079.\n[Zhao et al., 2021] Zhao, J., Khashabi, D., Khot, T., Sabharwal, A., and Chang, K.-W. (2021). Ethical-advice taker: Do language models understand natural language interventions?"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "llama",
      "Paper": "LLaMA: Open and Efficient Foundation Language Models",
      "AtlasYear": 2023,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2302.13971",
      "PdfSha256": "2E663675AE36AD12ADB2F5A05281BAC2747ECF8D23D92BEDD9F937A89FEE7136",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 603,
          "EndLine": 726,
          "PdfPages": [
            12,
            13,
            14,
            15,
            16
          ],
          "Text": "Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. Program synthesis with large language models.\nLalit R Bahl, Frederick Jelinek, and Robert L Mercer. 1983. A maximum likelihood approach to continuous speech recognition. IEEE transactions on pattern analysis and machine intelligence, pages 179– 190.\nYoshua Bengio, Réjean Ducharme, and Pascal Vincent. 2000. A neural probabilistic language model. Advances in neural information processing systems, 13.\nYonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artiﬁcial intelligence, pages 7432–7439.\nSid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745.\nThorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och, and Jeffrey Dean. 2007. Large language models in machine translation. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), pages 858–867, Prague, Czech Republic. Association for Computational Linguistics.\nPeter F Brown, John Cocke, Stephen A Della Pietra, Vincent J Della Pietra, Frederick Jelinek, John Lafferty, Robert L Mercer, and Paul S Roossin. 1990. A statistical approach to machine translation. Computational linguistics, 16(2):79–85.\nTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda\n\nAskell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.\nChristian Buck, Kenneth Heaﬁeld, and Bas Van Ooyen. 2014. N-gram counts and language models from the common crawl. In LREC, volume 2, page 4.\nCiprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005.\nMark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.\nAakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.\n\n\fHyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc Le, and Jason Wei. 2022. Scaling instruction-ﬁnetuned language models. arXiv preprint arXiv:2210.11416.\nChristopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difﬁculty of natural yes/no questions. arXiv preprint arXiv:1905.10044.\nPeter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.\nKarl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training veriﬁers to solve math word problems. arXiv preprint arXiv:2110.14168.\nZihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019. Transformer-xl: Attentive language models beyond a ﬁxed-length context. arXiv preprint arXiv:1901.02860.\nTri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashattention: Fast and memory-efﬁcient exact attention with io-awareness. arXiv preprint arXiv:2205.14135.\nJacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.\nJeffrey L Elman. 1990. Finding structure in time. Cognitive science, 14(2):179–211.\nDaniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wentau Yih, Luke Zettlemoyer, and Mike Lewis. 2022. Incoder: A generative model for code inﬁlling and synthesis. arXiv preprint arXiv:2204.05999.\nLeo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.\nLeo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPoﬁ, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff,\n\nJason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation.\n\nSamuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.\n\nAlex Graves. 2013. Generating sequences with\n\nrecurrent neural networks.\n\narXiv preprint\n\narXiv:1308.0850.\n\nKenneth Heaﬁeld, Ivan Pouzyrevsky, Jonathan H Clark, and Philipp Koehn. 2013. Scalable modiﬁed kneserney language model estimation. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 690–696.\n\nDan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.\n\nDan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874.\n\nJoel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou. 2017. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409.\n\nSepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9(8):1735–1780.\n\nJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training compute-optimal large language models.\n\nSrinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Dániel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, et al. 2022. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017.\n\nMandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551.\n\n\fRafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410.\nJared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.\nSlava Katz. 1987. Estimation of probabilities from sparse data for the language model component of a speech recognizer. IEEE transactions on acoustics, speech, and signal processing, 35(3):400–401.\nReinhard Kneser and Hermann Ney. 1995. Improved backing-off for m-gram language modeling. In 1995 international conference on acoustics, speech, and signal processing, volume 1, pages 181–184. IEEE.\nVijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2022. Reducing activation recomputation in large transformer models. arXiv preprint arXiv:2205.05198.\nTaku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226.\nKeita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Quantifying social biases in contextual word representations. In 1st ACL Workshop on Gender Bias for Natural Language Processing.\nTom Kwiatkowski, Jennimaria Palomaki, Olivia Redﬁeld, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.\nGuokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683.\nAitor Lewkowycz, Anders Johan Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Venkatesh Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving quantitative reasoning problems with language models. In Advances in Neural Information Processing Systems.\nOpher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham. 2021. Jurassic-1: Technical details and evaluation. White Paper. AI21 Labs, 1.\nStephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.\n\nIlya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.\n\nMatthew V Mahoney. 1999. Text compression as a test for artiﬁcial intelligence. AAAI/IAAI, 970.\n\nTodor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789.\n\nTomas Mikolov, Martin Karaﬁát, Lukas Burget, Jan Cernocky`, and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In Interspeech, pages 1045–1048. Makuhari.\n\nNikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In EMNLP 2020.\n\nErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474.\n\nLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.\n\nMarkus N Rabe and Charles Staats. 2021. attention does not need o(n2) memory.\npreprint arXiv:2112.05682.\n\nSelfarXiv\n\nAlec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving language understanding by generative pre-training.\n\nAlec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.\n\nJack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato,\n\n\fAngeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher.\nColin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a uniﬁed text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.\nJonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. 2019. A constructive prediction of the generalization error across scales. arXiv preprint arXiv:1909.12673.\nRachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender bias in coreference resolution. In NAACL-HLT 2018.\nKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106.\nMaarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728.\nTeven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic´, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176bparameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.\nRico Sennrich, Barry Haddow, and Alexandra Birch. 2015. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909.\nClaude E Shannon. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423.\nClaude E Shannon. 1951. Prediction and entropy of printed english. Bell system technical journal, 30(1):50–64.\nNoam Shazeer. 2020. Glu variants improve transformer. arXiv preprint arXiv:2002.05202.\n\nEmily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. arXiv preprint arXiv:1909.01326.\nMohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053.\nShaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model.\nJianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864.\nRomal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin HoffmanJohn, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022. Lamda: Language models for dialog applications.\nA. M. Turing. 1950. Computing Machinery and Intelligence. [Oxford University Press, Mind Association].\nAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30, pages 5998–6008.\nBen Wang and Aran Komatsuzaki. 2021. GPT-J6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/ mesh-transformer-jax.\nXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery,\n\n\fand Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models.\nJason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.\nGuillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020. CCNet: Extracting high quality monolingual datasets from web crawl data. In Language Resources and Evaluation Conference.\nCarole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. 2022. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4:795–813.\nRowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really ﬁnish your sentence? arXiv preprint arXiv:1905.07830.\nAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. 2022. Glm130b: An open bilingual pre-trained model.\nBiao Zhang and Rico Sennrich. 2019. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32.\nSusan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "gpt-4",
      "Paper": "GPT-4 Technical Report",
      "AtlasYear": 2023,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2303.08774",
      "PdfSha256": "C33A66DADCA2388D7B172D6293B00DC32B71110C6F38FAFE0D41112E61BE7774",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 765,
          "EndLine": 869,
          "PdfPages": [
            18,
            19,
            20,
            21,
            22,
            23
          ],
          "Text": "[1] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020.\n[2] Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.\n[3] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.\n[4] Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.\n[5] Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. Transformer-XL: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.\n[6] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019.\n[7] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.\n[8] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.\n[9] Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. arXiv preprint arXiv:1804.04235, 2018.\n[10] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.\n[11] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. NeurIPS, 2022.\n[12] Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. Large language models can self-improve. arXiv preprint arXiv:2210.11610, 2022.\n18\n\n\f[13] Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.\n[14] Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.\n[15] Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, et al. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701, 2020.\n[16] Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tensor Programs V: Tuning large neural networks via zero-shot hyperparameter transfer. arXiv preprint arXiv:2203.03466, 2022.\n[17] Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated Mixture-of-Experts layer. arXiv preprint arXiv:1701.06538, 2017.\n[18] Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. ST-MoE: Designing stable and transferable sparse expert models. arXiv preprint arXiv:2202.08906, 2022.\n[19] Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. TMLR, 2022.\n[20] Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyzdRiR9Y7.\n[21] Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. RoFormer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021.\n[22] Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems.\n[23] Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. PaLI: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022.\n[24] Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 billion parameter autoregressive language model, 2021.\n[25] Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow. If you use this software, please cite it using these metadata, 58, 2021.\n[26] Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic´, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176B-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.\n[27] Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. OPT: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.\n[28] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.\n[29] Alec Radford, Rafal Józefowicz, and Ilya Sutskever. Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444, 2017.\n19\n\n\f[30] Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291, 2019.\n[31] Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. arXiv preprint arXiv:2205.14135, 2022.\n[32] Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.\n[33] Markus N. Rabe and Charles Staats. Self-attention does not need o(n2) memory. arXiv preprint arXiv:2112.05682, 2021.\n[34] Scott Gray, Alec Radford, and Diederik P. Kingma. Gpu kernels for block-sparse weights, 2017. URL https://cdn.openai.com/blocksparse/blocksparsepaper.pdf.\n[35] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021.\n[36] Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning AI with shared human values. Proceedings of the International Conference on Learning Representations (ICLR), 2021.\n[37] Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.\n[38] Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018.\n[39] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017.\n[40] Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30, 2017.\n[41] Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409, 2017.\n[42] Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso. The computational limits of deep learning. arXiv preprint arXiv:2007.05558, 2020.\n[43] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. 2021.\n[44] Ian McKenzie, Alexander Lyzhov, Alicia Parrish, Ameya Prabhu, Aaron Mueller, Najoung Kim, Sam Bowman, and Ethan Perez. The Inverse Scaling Prize, 2022. URL https://github. com/inverse-scaling/prize.\n[45] Jason Wei, Najoung Kim, Yi Tay, and Quoc V. Le. Inverse scaling can become U-shaped. arXiv preprint arXiv:2211.02011, 2022.\n[46] Ian McKenzie, Alexander Lyzhov, Alicia Parrish, Ameya Prabhu, Aaron Mueller, Najoung Kim, Sam Bowman, and Ethan Perez. Inverse Scaling Prize: First round winners, 2022. URL https://irmckenzie.co.uk/round1.\n20\n\n\f[47] Greg Brockman, Peter Welinder, Mira Murati, and OpenAI. OpenAI: OpenAI API, 2020. URL https://openai.com/blog/openai-api.\n[48] Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.\n[49] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.\n[50] Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q Tran, David R So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, et al. Transcending scaling laws with 0.1% extra compute. arXiv preprint arXiv:2210.11399, 2022.\n[51] Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.\n[52] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1472. URL https://aclanthology.org/P19-1472.\n[53] Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. Adversarial training for large neural language models. arXiv preprint arXiv:2004.08994, 2020.\n[54] Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. ArXiv, abs/1803.05457, 2018.\n[55] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. Selfconsistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.\n[56] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An adversarial Winograd schema challenge at scale. arXiv preprint arXiv:1907.10641, 2019.\n[57] Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. CodeT: Code generation with generated tests. arXiv preprint arXiv:2207.10397, 2022.\n[58] Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1246. URL https://aclanthology. org/N19-1246.\n[59] Kunlong Chen, Weidi Xu, Xingyi Cheng, Zou Xiaochuan, Yuyu Zhang, Le Song, Taifeng Wang, Yuan Qi, and Wei Chu. Question directed graph attention network for numerical reasoning over text. arXiv preprint arXiv:2009.07448, 2020.\n[60] Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.\n[61] Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. Solving quantitative reasoning problems with language models. arXiv preprint arXiv:2206.14858, 2022.\n21\n\n\f[62] Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process- and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022.\n[63] Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.\n[64] OpenAI. OpenAI: Introducing ChatGPT, 2022. URL https://openai.com/blog/chatgpt.\n[65] OpenAI. OpenAI: GPT-4, 2023. URL https://openai.com/research/gpt-4.\n[66] Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-long.229.\n[67] Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.\n[68] OpenAI. OpenAI: How should AI systems behave, and who should decide?, 2023. URL https://openai.com/blog/how-should-ai-systems-behave.\n[69] Jan Leike, John Schulman, and Jeffrey Wu. OpenAI: Our approach to alignment research, 2022. URL https://openai.com/blog/our-approach-to-alignment-research.\n[70] Joseph Carlsmith. Is power-seeking AI an existential risk? ArXiv, abs/2206.13353, 2022.\n[71] Amelia Glaese, Nat McAleese, Maja Tre˛bacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Sonˇa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.\n[72] Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. arXiv preprint arXiv:2202.03286, 2022.\n[73] Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462, 2020.\n[74] Dora Seigel. How do you calculate SAT score? raw and scaled, 1 2020. URL https: //blog.prepscholar.com/how-to-calculate-sat-score.\n[75] The Albert blog. URL https://www.albert.io/blog/.\n[76] Mathematical Association of America. AMC statistics, 2023. URL http://amc-reg.maa. org/Reports/GeneralReports.aspx.\n[77] Halle Edwards. SAT percentiles and score rankings, 2022. URL https://blog. prepscholar.com/sat-percentiles-and-score-rankings.\n[78] College Board. Understanding SAT scores, 2022. URL https://satsuite.collegeboard. org/media/pdf/understanding-sat-scores.pdf.\n[79] College Board. AP score distributions by subject, 2022. URL https://apcentral. collegeboard.org/media/pdf/ap-score-distributions-by-subject-2022.pdf.\n22\n\n\f[80] Center for Excellence in Education. 2020 USABO Semifinal exam score distribution, 2022. URL https://www.usabo-trc.org/sites/default/files/allfiles/2020% 20USABO%20Semifinal%20Exam%20Histogram.pdf.\n\n[81] Chris Swimmer. GRE score percentiles – what does your score mean for you? (2021 update), 4 2021. URL https://magoosh.com/gre/gre-score-percentiles/.\n\n[82] John B. Nici. AP Art History: 5 Practice Tests + Comprehensive Review + Online Practice. Barron’s Test Prep. Barron’s Educational Series, 2020. ISBN 9781506260501.\n\n[83] ETS. GRE sample issue task, 2022. sample-issue-task.pdf.\n\nURL https://www.ets.org/pdfs/gre/\n\n[84] Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pages 220– 229, January 2019. doi: 10.1145/3287560.3287596.\n\n[85] Nekesha Green, Chavez Procope, Adeel Cheema, and Adekunle Adediji. System Cards, a new resource for understanding how AI systems work. https://ai.facebook.com/blog/system-cards-anew-resource-for-understanding-how-ai-systems-work/, February 2022.\n\n23"
        },
        {
          "Section": "Reference section 2",
          "StartLine": 2043,
          "EndLine": 2170,
          "PdfPages": [
            71,
            72,
            73,
            74,
            75,
            76,
            77,
            78
          ],
          "Text": "[1] A. Tamkin, M. Brundage, J. Clark, and D. Ganguli, “Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models,” Feb. 2021.\n[2] “Introducing the new Bing.” https://www.bing.com/new.\n[3] J. Hilton, R. Nakano, S. Balaji, and J. Schulman, “WebGPT: Improving the factual accuracy of language models through web browsing.” https://openai.com/research/webgpt, Dec. 2021.\n[4] “ACT-1: Transformer for Actions – Adept.” https://www.adept.ai/blog/act-1.\n[5] M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba, “Evaluating Large Language Models Trained on Code,” July 2021.\n[6] L. Weidinger, J. Mellor, M. Rauh, C. Griﬃn, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. Isaac, S. Legassick, G. Irving, and I. Gabriel, “Ethical and social risks of harm from Language Models,” Dec. 2021.\n[7] I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, G. Krueger, J. W. Kim, S. Kreps, M. McCain, A. Newhouse, J. Blazakis, K. McGuﬃe, and J. Wang, “Release Strategies and the Social Impacts of Language Models,” Nov. 2019.\n[8] A. Radford, “Improving language understanding with unsupervised learning.” https://openai.com/research/language-unsupervised, June 2018.\n[9] A. Radford, J. Wu, D. Amodei, D. Amodei, J. Clark, M. Brundage, I. Sutskever, A. Askell, D. Lansky, D. Hernandez, and D. Luan, “Better language models and their implications.” https://openai.com/research/better-language-models, Feb. 2019.\n[10] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language Models are Few-Shot Learners,” July 2020.\n[11] S. Altman, “Planning for AGI and beyond.” https://openai.com/blog/planning-for-agi-andbeyond, Feb. 2023.\n[12] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” Mar. 2022.\n71\n\n\f[13] P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Feb. 2023.\n[14] M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru, “Model Cards for Model Reporting,” in Proceedings of the Conference on Fairness, Accountability, and Transparency, pp. 220–229, Jan. 2019.\n[15] N. Green, C. Procope, A. Cheema, and A. Adediji, “System Cards, a new resource for understanding how AI systems work.” https://ai.facebook.com/blog/system-cards-a-new-resourcefor-understanding-how-ai-systems-work/, Feb. 2022.\n[16] “DALL·E 2 Preview - Risks and Limitations.” OpenAI, Apr. 2022.\n[17] J. Sandbrink, H. Hobbs, J. Swett, A. Dafoe, and A. Sandberg, “Diﬀerential Technology Development: A Responsible Innovation Principle for Navigating Technology Risks,” Sept. 2022.\n[18] Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatﬁeld-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan, “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback,” Apr. 2022.\n[19] E. Perez, S. Ringer, K. Lukošiu¯te˙, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, A. Jones, A. Chen, B. Mann, B. Israel, B. Seethor, C. McKinnon, C. Olah, D. Yan, D. Amodei, D. Amodei, D. Drain, D. Li, E. Tran-Johnson, G. Khundadze, J. Kernion, J. Landis, J. Kerr, J. Mueller, J. Hyun, J. Landau, K. Ndousse, L. Goldberg, L. Lovitt, M. Lucas, M. Sellitto, M. Zhang, N. Kingsland, N. Elhage, N. Joseph, N. Mercado, N. DasSarma, O. Rausch, R. Larson, S. McCandlish, S. Johnston, S. Kravec, S. E. Showk, T. Lanham, T. Telleen-Lawton, T. Brown, T. Henighan, T. Hume, Y. Bai, Z. Hatﬁeld-Dodds, J. Clark, S. R. Bowman, A. Askell, R. Grosse, D. Hernandez, D. Ganguli, E. Hubinger, N. Schiefer, and J. Kaplan, “Discovering Language Model Behaviors with Model-Written Evaluations,” Dec. 2022.\n[20] B. P. Kehoe, Zen and the Art of the Internet. Project Gutenberg, June 1992.\n[21] M. Brundage, K. Mayer, T. Eloundou, S. Agarwal, S. Adler, G. Krueger, J. Leike, and P. Mishkin, “Lessons learned on language model safety and misuse.” https://openai.com/research/language-model-safety-and-misuse, Mar. 2022.\n[22] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” 2019.\n[23] G. C. Bowker and S. L. Star, Sorting Things Out. MIT Press, Aug. 2000.\n[24] L. Weidinger, J. Uesato, M. Rauh, C. Griﬃn, P.-S. Huang, J. Mellor, A. Glaese, M. Cheng, B. Balle, A. Kasirzadeh, C. Biles, S. Brown, Z. Kenton, W. Hawkins, T. Stepleton, A. Birhane, L. A. Hendricks, L. Rimell, W. Isaac, J. Haas, S. Legassick, G. Irving, and I. Gabriel, “Taxonomy of Risks posed by Language Models,” in 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, (New York, NY, USA), pp. 214–229, Association for Computing Machinery, June 2022.\n72\n\n\f[25] I. Solaiman and C. Dennison, “Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets,” Nov. 2021.\n[26] H. Khlaaf, “Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems,” Trail of Bits, 2023.\n[27] M. Brundage, S. Avin, J. Wang, H. Belﬁeld, G. Krueger, G. Hadﬁeld, H. Khlaaf, J. Yang, H. Toner, R. Fong, T. Maharaj, P. W. Koh, S. Hooker, J. Leung, A. Trask, E. Bluemke, J. Lebensold, C. O’Keefe, M. Koren, T. Ryﬀel, J. B. Rubinovitz, T. Besiroglu, F. Carugati, J. Clark, P. Eckersley, S. de Haas, M. Johnson, B. Laurie, A. Ingerman, I. Krawczuk, A. Askell, R. Cammarota, A. Lohn, D. Krueger, C. Stix, P. Henderson, L. Graham, C. Prunkl, B. Martin, E. Seger, N. Zilberman, S. Ó. hÉigeartaigh, F. Kroeger, G. Sastry, R. Kagan, A. Weller, B. Tse, E. Barnes, A. Dafoe, P. Scharre, A. Herbert-Voss, M. Rasser, S. Sodhani, C. Flynn, T. K. Gilbert, L. Dyer, S. Khan, Y. Bengio, and M. Anderljung, “Toward Trustworthy AI Development: Mechanisms for Supporting Veriﬁable Claims,” Apr. 2020.\n[28] D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, A. Jones, S. Bowman, A. Chen, T. Conerly, N. DasSarma, D. Drain, N. Elhage, S. El-Showk, S. Fort, Z. Hatﬁeld-Dodds, T. Henighan, D. Hernandez, T. Hume, J. Jacobson, S. Johnston, S. Kravec, C. Olsson, S. Ringer, E. Tran-Johnson, D. Amodei, T. Brown, N. Joseph, S. McCandlish, C. Olah, J. Kaplan, and J. Clark, “Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned,” Nov. 2022.\n[29] E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving, “Red Teaming Language Models with Language Models,” Feb. 2022.\n[30] H. Khlaaf, P. Mishkin, J. Achiam, G. Krueger, and M. Brundage, “A Hazard Analysis Framework for Code Synthesis Large Language Models,” July 2022.\n[31] J. Maynez, S. Narayan, B. Bohnet, and R. McDonald, “On Faithfulness and Factuality in Abstractive Summarization,” May 2020.\n[32] S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring How Models Mimic Human Falsehoods,” May 2022.\n[33] J. A. Goldstein, G. Sastry, M. Musser, R. DiResta, M. Gentzel, and K. Sedova, “Forecasting potential misuses of language models for disinformation campaigns and how to reduce risk.” https://openai.com/research/forecasting-misuse, Jan. 2023.\n[34] O. Evans, O. Cotton-Barratt, L. Finnveden, A. Bales, A. Balwit, P. Wills, L. Righetti, and W. Saunders, “Truthful AI: Developing and governing AI that does not lie,” Oct. 2021.\n[35] A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein, “Detoxifying Language Models Risks Marginalizing Minority Voices,” Apr. 2021.\n[36] L. Dixon, J. Li, J. Sorensen, N. Thain, and L. Vasserman, “Measuring and Mitigating Unintended Bias in Text Classiﬁcation,” in Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’18, (New York, NY, USA), pp. 67–73, Association for Computing Machinery, Dec. 2018.\n[37] T. Markov, C. Zhang, S. Agarwal, T. Eloundou, T. Lee, S. Adler, A. Jiang, and L. Weng, “A Holistic Approach to Undesired Content Detection in the Real World,” Feb. 2023.\n73\n\n\f[38] OpenAI, “How should AI systems behave, and who should decide?.” https://openai.com/blog/how-should-ai-systems-behave, Feb. 2023.\n[39] M. Rauh, J. Mellor, J. Uesato, P.-S. Huang, J. Welbl, L. Weidinger, S. Dathathri, A. Glaese, G. Irving, I. Gabriel, W. Isaac, and L. A. Hendricks, “Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models,” Oct. 2022.\n[40] S. L. Blodgett, S. Barocas, H. Daumé III, and H. Wallach, “Language (Technology) is Power: A Critical Survey of \"Bias\" in NLP.” https://arxiv.org/abs/2005.14050v2, May 2020.\n[41] S. Dev, E. Sheng, J. Zhao, A. Amstutz, J. Sun, Y. Hou, M. Sanseverino, J. Kim, A. Nishi, N. Peng, and K.-W. Chang, “On Measures of Biases and Harms in NLP,” in Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, (Online only), pp. 246–267, Association for Computational Linguistics, Nov. 2022.\n[42] T. Bolukbasi, K.-W. Chang, J. Zou, V. Saligrama, and A. Kalai, “Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings,” July 2016.\n[43] H. Gonen and Y. Goldberg, “Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), (Minneapolis, Minnesota), pp. 609–614, Association for Computational Linguistics, June 2019.\n[44] K. Webster, M. Recasens, V. Axelrod, and J. Baldridge, “Mind the GAP: A Balanced Corpus of Gendered Ambiguous Pronouns,” Oct. 2018.\n[45] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? ,” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, (Virtual Event Canada), pp. 610–623, ACM, Mar. 2021.\n[46] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang, “On the Opportunities and Risks of Foundation Models,” Aug. 2021.\n[47] S. U. Noble, Algorithms of Oppression. NYU Press, Feb. 2018.\n[48] R. Richardson, J. Schultz, and K. Crawford, “Dirty Data, Bad Predictions: How Civil Rights Violations Impact Police Data, Predictive Policing Systems, and Justice,” Feb. 2019.\n74\n\n\f[49] W. MacAskill, What We Owe The Future. Basic Books, Aug. 2022.\n[50] OpenAI, “GPT-2: 1.5B release.” https://openai.com/research/gpt-2-1-5b-release, Nov. 2019.\n[51] S. Kreps, R. M. McCain, and M. Brundage, “All the News That’s Fit to Fabricate: AIGenerated Text as a Tool of Media Misinformation,” Journal of Experimental Political Science, vol. 9, no. 1, pp. 104–117, 2022/ed.\n[52] B. Buchanan, A. Lohn, M. Musser, and K. Sedova, “Truth, Lies, and Automation,” tech. rep., Center for Security and Emerging Technology, May 2021.\n[53] A. Myers, “AI’s Powers of Political Persuasion.” https://hai.stanford.edu/news/ais-powerspolitical-persuasion, Feb. 2023.\n[54] H. Bai, J. Voelkel, J. Eichstaedt, and R. Willer, “Artiﬁcial intelligence can persuade humans on political issues,” 2023.\n[55] E. Horvitz, “On the Horizon: Interactive and Compositional Deepfakes,” in INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION, pp. 653–661, Nov. 2022.\n[56] R. Chesney and D. K. Citron, “Deep Fakes: A Looming Challenge for Privacy, Democracy, and National Security,” July 2018.\n[57] U.S. Department of Commerce, “Dual use export licenses,” March 13 2023. accessed 2023-03-13.\n[58] NATO, “Arms control, disarmament and non-proliferation in nato,” February 27 2023. accessed 2023-02-27.\n[59] N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raﬀel, “Extracting Training Data from Large Language Models,” June 2021.\n[60] N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, “Quantifying Memorization Across Neural Language Models,” Mar. 2023.\n[61] D. Ganguli, D. Hernandez, L. Lovitt, N. DasSarma, T. Henighan, A. Jones, N. Joseph, J. Kernion, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, N. Elhage, S. E. Showk, S. Fort, Z. Hatﬁeld-Dodds, S. Johnston, S. Kravec, N. Nanda, K. Ndousse, C. Olsson, D. Amodei, D. Amodei, T. Brown, J. Kaplan, S. McCandlish, C. Olah, and J. Clark, “Predictability and Surprise in Large Generative Models,” in 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 1747–1764, June 2022.\n[62] J. Wei, Y. Tay, R. Bommasani, C. Raﬀel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent Abilities of Large Language Models,” Oct. 2022.\n[63] R. Ngo, L. Chan, and S. Mindermann, “The alignment problem from a deep learning perspective,” Feb. 2023.\n[64] N. Bostrom, Superintelligence: Paths, Dangers, Strategies. United Kingdom: Oxford University Press, Sept. 2014.\n75\n\n\f[65] A. Chan, R. Salganik, A. Markelius, C. Pang, N. Rajkumar, D. Krasheninnikov, L. Langosco, Z. He, Y. Duan, M. Carroll, M. Lin, A. Mayhew, K. Collins, M. Molamohammadi, J. Burden, W. Zhao, S. Rismani, K. Voudouris, U. Bhatt, A. Weller, D. Krueger, and T. Maharaj, “Harms from Increasingly Agentic Algorithmic Systems,” Feb. 2023.\n[66] J. Andreas, “Language Models as Agent Models,” Dec. 2022.\n[67] J. Steinhardt, “Emergent Deception and Emergent Optimization.” https://boundedregret.ghost.io/emergent-deception-optimization/, Feb. 2023.\n[68] S. M. Omohundro, “The Basic AI Drives,” in Proceedings of the 2008 Conference on Artiﬁcial General Intelligence 2008, (NLD), pp. 483–492, IOS Press, June 2008.\n[69] N. Bostrom, “The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artiﬁcial Agents,” Minds and Machines, vol. 22, pp. 71–85, May 2012.\n[70] A. M. Turner, L. Smith, R. Shah, A. Critch, and P. Tadepalli, “Optimal Policies Tend to Seek Power,” Jan. 2023.\n[71] A. M. Turner and P. Tadepalli, “Parametrically Retargetable Decision-Makers Tend To Seek Power,” Oct. 2022.\n[72] V. Krakovna and janos, “Power-seeking can be probable and predictive for trained agents,” Mar. 2023.\n[73] S. Russell, Human Compatible: Artiﬁcial Intelligence and the Problem of Control. Cham: Springer International Publishing, 2022.\n[74] J. Carlsmith, “Is Power-Seeking AI an Existential Risk?,” June 2022.\n[75] Alignment Research Center, “Update on arc’s recent eval eﬀorts,” March 2023 2023. accessed 2023-03-17.\n[76] E. Karpas, O. Abend, Y. Belinkov, B. Lenz, O. Lieber, N. Ratner, Y. Shoham, H. Bata, Y. Levine, K. Leyton-Brown, D. Muhlgay, N. Rozen, E. Schwartz, G. Shachaf, S. ShalevShwartz, A. Shashua, and M. Tenenholtz, “MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning,” May 2022.\n[77] T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language Models Can Teach Themselves to Use Tools,” Feb. 2023.\n[78] G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celikyilmaz, E. Grave, Y. LeCun, and T. Scialom, “Augmented Language Models: A Survey,” Feb. 2023.\n[79] A. Parisi, Y. Zhao, and N. Fiedel, “TALM: Tool Augmented Language Models,” May 2022.\n[80] D. Weininger, “Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules,” Journal of chemical information and computer sciences, vol. 28, no. 1, pp. 31–36, 1988.\n[81] E. Calvano, G. Calzolari, V. Denicolò, and S. Pastorello, “Artiﬁcial Intelligence, Algorithmic Pricing and Collusion,” Apr. 2019.\n76\n\n\f[82] D. Krueger, T. Maharaj, and J. Leike, “Hidden Incentives for Auto-Induced Distributional Shift,” Sept. 2020.\n[83] S. J. DeCanio, “Robots and humans – complements or substitutes?,” Journal of Macroeconomics, vol. 49, pp. 280–291, Sept. 2016.\n[84] A. Korinek and J. E. Stiglitz, “Artiﬁcial Intelligence and Its Implications for Income Distribution and Unemployment,” in The Economics of Artiﬁcial Intelligence: An Agenda, pp. 349–390, University of Chicago Press, Jan. 2018.\n[85] J. H. Choi, K. E. Hickman, A. Monahan, and D. Schwarcz, “ChatGPT Goes to Law School,” Jan. 2023.\n[86] L. R. Raymond, E. Brynjolfsson, and D. Li, “Augmented intelligence: The eﬀects of ai on productivity and work practices,” Sep 2022.\n[87] E. van Inwegen, Z. Munyikwa, and J. J. Horton, “Algorithmic Writing Assistance on Jobseekers’ Resumes Increases Hires,” Jan. 2023.\n[88] A. Ziegler, E. Kalliamvakou, S. Simister, G. Sittampalam, A. Li, A. Rice, D. Rifkin, and E. Aftandilian, “Productivity Assessment of Neural Code Completion,” May 2022.\n[89] S. Noy and W. Zhang, “Experimental evidence on the productivity eﬀects of generative artiﬁcial intelligence,” Available at SSRN 4375283, 2023.\n[90] S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer, “The impact of ai on developer productivity: Evidence from github copilot,” arXiv preprint arXiv:2302.06590, 2023.\n[91] D. Acemoglu and P. Restrepo, “Demographics and Automation,” The Review of Economic Studies, vol. 89, pp. 1–44, Jan. 2022.\n[92] Partnership on AI, “AI and Job Quality,” tech. rep., Partnership on AI, Sept. 2022.\n[93] “OpenAI Charter.” https://openai.com/charter, Apr. 2018.\n[94] S. Armstrong, N. Bostrom, and C. Shulman, “Racing to the precipice: A model of artiﬁcial intelligence development,” Technical 2013-1, Future of Humanity Institute, Oct. 2013.\n[95] P. E. Tetlock and D. Gardner, Superforecasting: The Art and Science of Prediction. Crown, Sept. 2015.\n[96] S. Passi and M. Vorvoreanu, “Overreliance on AI Literature Review,” tech. rep., AI Ethics and Eﬀects in Engineering and Research, June 2022.\n[97] PAI, “Data enrichment sourcing guidelines,” November 2022 2022. accessed 2023-03-13.\n[98] PAI, “Responsible sourcing of data enrichment services,” June 2021 2021. accessed 2023-03-13.\n[99] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” Aug. 2017.\n77\n\n\f[100]\n\nA. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, L. Campbell-Gillingham, J. Uesato, P.-S. Huang, R. Comanescu, F. Yang, A. See, S. Dathathri, R. Greig, C. Chen, D. Fritz, J. S. Elias, R. Green, S. Mokrá, N. Fernando, B. Wu, R. Foley, S. Young, I. Gabriel, W. Isaac, J. Mellor, D. Hassabis, K. Kavukcuoglu, L. A. Hendricks, and G. Irving, “Improving alignment of dialogue agents via targeted human judgements,” Sept. 2022.\n\n[101]\n\nY. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatﬁeld-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan, “Constitutional AI: Harmlessness from AI Feedback,” Dec. 2022.\n\n[102] S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models,” Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 3356–3369, 2020.\n\n[103] OpenAI, “Introducing chatgpt,” November 2022 2020. accessed 2023-03-13.\n\n[104] OpenAI, “Openai api,” June 2020 2020. accessed 2023-03-13.\n\n[105] T. Davidson, D. Bhattacharya, and I. Weber, “Racial Bias in Hate Speech and Abusive Language Detection Datasets,” in Proceedings of the Third Workshop on Abusive Language Online, (Florence, Italy), pp. 25–35, Association for Computational Linguistics, Aug. 2019."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified.",
        "Includes separate bibliographies for the technical report and appended system card."
      ]
    },
    {
      "Slug": "mistral-7b",
      "Paper": "Mistral 7B",
      "AtlasYear": 2023,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2310.06825",
      "PdfSha256": "DFBAC4E7035344B305C947481F2E7E8A02F7A24A563917EB6E47F6591D14C5AE",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 167,
          "EndLine": 200,
          "PdfPages": [
            8,
            9,
            10
          ],
          "Text": "[1] Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. arXiv preprint arXiv:2305.13245, 2023.\n[2] Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.\n[3] Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.\n[4] Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, 2020.\n[5] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.\n[6] Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.\n[7] Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. Quac: Question answering in context. arXiv preprint arXiv:1808.07036, 2018.\n[8] Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.\n[9] Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.\n[10] Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.\n[11] Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems, 2022.\n[12] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.\n[13] Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.\n[14] Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karén Simonyan, Erich Elsen, Oriol Vinyals, Jack Rae, and Laurent Sifre. An empirical analysis of compute-optimal large language model training. In Advances in Neural Information Processing Systems, volume 35, 2022.\n[15] Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017.\n[16] Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466, 2019.\n8\n\n\f[17] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.\n[18] Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, and Daniel Haziza. xformers: A modular and hackable transformer modelling library. https://github.com/ facebookresearch/xformers, 2022.\n[19] Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.\n[20] Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.\n[21] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.\n[22] Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728, 2019.\n[23] Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, , and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.\n[24] Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937, 2018.\n[25] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.\n[26] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.\n[27] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.\n[28] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.\n[29] Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364, 2023.\n9"
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "mixtral",
      "Paper": "Mixtral of Experts",
      "AtlasYear": 2024,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2401.04088",
      "PdfSha256": "F8BBF0E9D979B7A8CE7BE65119266545A229A85B57E077D8BD048E458BB642DA",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 302,
          "EndLine": 340,
          "PdfPages": [
            9,
            10,
            11
          ],
          "Text": "[1] Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.\n[2] Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics. arXiv preprint arXiv:2310.10631, 2023.\n[3] Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, pages 7432–7439, 2020.\n[4] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.\n[5] Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. Quac: Question answering in context. arXiv preprint arXiv:1808.07036, 2018.\n[6] Aidan Clark, Diego De Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al. Unified scaling laws for routed language models. In International Conference on Machine Learning, pages 4057–4086. PMLR, 2022.\n[7] Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.\n[8] Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.\n[9] Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.\n[10] Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. Bold: Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 862–872, 2021.\n[11] Artyom Eliseev and Denis Mazur. Fast inference of mixture-of-experts language models with offloading. arXiv preprint arXiv:2312.17238, 2023.\n[12] William Fedus, Jeff Dean, and Barret Zoph. A review of sparse expert models in deep learning. arXiv preprint arXiv:2209.01667, 2022.\n[13] Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia. Megablocks: Efficient sparse training with mixture-of-experts. arXiv preprint arXiv:2211.15841, 2022.\n[14] Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.\n[15] Hussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Yihua Chen, Rahul Mazumder, Lichan Hong, and Ed Chi. Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning. Advances in Neural Information Processing Systems, 34:29335–29347, 2021.\n9\n\n\f[16] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.\n[17] Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.\n[18] Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.\n[19] Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551, 2017.\n[20] Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, pages 453–466, 2019.\n[21] Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.\n[22] Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.\n[23] Amirkeivan Mohtashami and Martin Jaggi. Landmark attention: Random-access infinite context length for transformers. arXiv preprint arXiv:2305.16300, 2023.\n[24] Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. Bbq: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193, 2021.\n[25] Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.\n[26] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, pages 99–106, 2021.\n[27] Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728, 2019.\n[28] Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.\n[29] Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, , and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.\n[30] Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937, 2018.\n[31] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.\n[32] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.\n[33] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.\n10\n\n\f[34] Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364, 2023.\n[35] Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew M Dai, Quoc V Le, James Laudon, et al. Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems, 35:7103–7114, 2022."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "gemini-1-5",
      "Paper": "Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context",
      "AtlasYear": 2024,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2403.05530",
      "PdfSha256": "ABCD3B8E4B54721FF2D2323204F7B2AD9212B0FF6C8ED77725657526A5418B13",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 2904,
          "EndLine": 3319,
          "PdfPages": [
            75,
            76,
            77,
            78,
            79,
            80,
            81,
            82,
            83,
            84,
            85,
            86,
            87,
            88,
            89,
            90,
            91
          ],
          "Text": "Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Stephanie Chan, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle. Many-shot in-context learning. CoRR, abs/2404.11018, 2024a.\nRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations, 2024b.\nJoshua Ainslie, Tao Lei, Michiel de Jong, Santiago Ontañón, Siddhartha Brahma, Yury Zemlyanskiy, David Uthus, Mandy Guo, James Lee-Thorp, Yi Tay, et al. Colt5: Faster long-range transformers with conditional computation. arXiv preprint arXiv:2303.09752, 2023.\nMaksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safetyaligned llms with simple adaptive attacks. arXiv preprint arXiv:2404.02151, 2024.\nRohan Anil, Gabriel Pereyra, Alexandre Passos, Robert Ormandi, George E. Dahl, and Geoffrey E. Hinton. Large scale distributed neural network training through online distillation, 2018.\nRohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023a.\nRohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. PaLM 2 Technical Report, 2023b.\nAnthropic. Model Card and Evaluations for Claude Models, 2023a.\nAnthropic. Long context prompting for Claude 2.1, 2023b.\nRosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. Common voice: A massivelymultilingual speech corpus. arXiv preprint arXiv:1912.06670, 2019.\nYuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073, 2022.\nIvana Balažević, Yuge Shi, Pinelopi Papalampidi, Rahma Chaabouni, Skanda Koppula, and Olivier J Hénaff. Memory consolidation enables long-context video understanding. arXiv preprint arXiv:2402.05861, 2024.\nSuzanna Becker and Yann LeCun. Improving the convergence of back-propagation learning\nwith second-order methods. 1989. URL https://api.semanticscholar.org/CorpusID: 59695337.\n75\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nYoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation, 2013.\nAmanda Bertsch, Uri Alon, Graham Neubig, and Matthew R Gormley. Unlimiformer: Long-range transformers with unlimited length input. arXiv preprint arXiv:2305.01625, 2023.\nAmanda Bertsch, Maor Ivgi, Uri Alon, Jonathan Berant, Matthew R Gormley, and Graham Neubig. In-context learning with long-context models: An in-depth exploration. arXiv preprint arXiv:2405.00200, 2024.\nLucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov. Knowledge distillation: A good teacher is patient and consistent, 2021.\nStella Biderman, USVSN PRASHANTH, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raff. Emergent and predictable memorization in large language models. Advances in Neural Information Processing Systems, 36, 2024.\nSteven Bird. Decolonising speech and language technology. In Donia Scott, Nuria Bel, and Chengqing Zong, editors, Proceedings of the 28th International Conference on Computational Linguistics, pages 3504–3519, Barcelona, Spain (Online), December 2020. International Committee on Computational\nLinguistics. doi: 10.18653/v1/2020.coling-main.313. URL https://aclanthology.org/2020. coling-main.313.\nAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluis Gomez, Marçal Rusiñol, C.V. Jawahar, Ernest Valveny, and Dimosthenis Karatzas. Scene text visual question answering. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 4290–4300, 2019. doi: 10.1109/ICCV.2019.00439.\nBernd Bohnet, Kevin Swersky, Rosanne Liu, Pranjal Awasthi, Azade Nova, Javier Snaider, Hanie Sedghi, Aaron T Parisi, Michael Collins, Angeliki Lazaridou, Orhan Firat, and Noah Fiedel. Longspan question-answering: Automatic question generation and qa-system ranking via side-by-side evaluation, 2024.\nJames Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX:\ncomposable transformations of Python+NumPy programs, 2018. URL http://github.com/ google/jax.\nRalph A. Bradley and Milton E. Terry. The rank analysis of incomplete block designs — I. The method of paired comparisons. Biometrika, 39:324–345, 1952.\nThorsten Brants, Ashok Popat, Peng Xu, Franz Och, and Jeffrey Dean. Large language models in machine translation. pages 858–867, 01 2007.\nAndrei Broder. A taxonomy of web search. SIGIR Forum, 36(2):3–10, sep 2002. ISSN 0163-5840.\ndoi: 10.1145/792550.792552. URL https://doi.org/10.1145/792550.792552.\nTom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel HerbertVoss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages\n1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_ files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.\n76\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nCristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil. Model compression. In Knowl-\nedge Discovery and Data Mining, 2006. URL https://api.semanticscholar.org/CorpusID: 11253972.\nAydar Bulatov, Yury Kuratov, and Mikhail Burtsev. Recurrent memory transformer. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 11079–11091. Curran Associates,\nInc., 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/ 47e288629a6996a17ce50b90a056a0e1-Paper-Conference.pdf.\nAydar Bulatov, Yuri Kuratov, and Mikhail S Burtsev. Scaling transformer to 1m tokens and beyond with rmt. arXiv preprint arXiv:2304.11062, 2023.\nJoy Buolamwini and Timnit Gebru. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency, pages 77–91. PMLR, 2018.\nNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646,\n2022. URL https://arxiv.org/abs/2202.07646.\nNicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt. Are aligned neural networks adversarially aligned?, 2024.\nBen Carterette, Paul Clough, Mark Hall, Evangelos Kanoulas, and Mark Sanderson. Evaluating retrieval over sessions: The TREC session track 2011-2014. In Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’16, page 685–688, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 9781450340694. doi:\n10.1145/2911451.2914675. URL https://doi.org/10.1145/2911451.2914675.\nPatrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong. Jailbreakbench: An open robustness benchmark for jailbreaking large language models, 2024.\nMark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. arXiv\npreprint arXiv:2107.03374, 2021. URL https://arxiv.org/abs/2107.03374.\nStanley F Chen and Joshua Goodman. An empirical study of smoothing techniques for language modeling. Computer Speech & Language, 13(4):359–394, 1999.\nYizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner. DiverseVul: A new vulnerable source code dataset for deep learning based vulnerability detection. In International\n77\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nSymposium on Research in Attacks, Intrusions and Defenses, pages 654–668, April 2023a. URL\nhttps://arxiv.org/abs/2304.00409.\nYukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. Longlora: Efficient fine-tuning of long-context large language models. arXiv preprint arXiv:2309.12307,\n2023b. URL https://arxiv.org/abs/2309.12307.\nAakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023a.\nAakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. PaLM: Scaling Language Modeling with Pathways. Journal of Machine Learning Research, 24(240):1–113, 2023b.\nURL http://jmlr.org/papers/v24/22-1144.html.\nAidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Hennigan, Matthew Johnson, Katie Millican, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Jack Rae, Erich Elsen, Koray Kavukcuoglu, and Karen Simonyan. Unified scaling laws for routed language models, 2022.\nURL https://arxiv.org/abs/2202.01169.\nPeter Clark, Oyvind Tafjord, and Kyle Richardson. Transformers as soft reasoners over language.\nIJCAI, 2020. URL https://www.ijcai.org/proceedings/2020/0537.pdf.\nCharles LA Clarke, Saira Rizvi, Mark D Smucker, Maria Maistro, and Guido Zuccon. Overview of the TREC 2020 health misinformation track. In Proceedings of the Twenty-Ninth Text REtrieval Conference\n(TREC 2020), 2020. URL https://trec.nist.gov/pubs/trec29/papers/OVERVIEW.HM. pdf.\nKarl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint\narXiv:2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.\nKevyn Collins-Thompson, Craig Macdonald, Paul N. Bennett, Fernando Diaz, and Ellen M. Voorhees. TREC 2014 web track overview. In Proceedings of the Twenty-Third Text REtrieval Conference (TREC\n2014), 2014. URL https://trec.nist.gov/pubs/trec23/papers/overview-web.pdf.\nAlexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. Fleurs: Few-shot learning evaluation of universal representations of speech. In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805. IEEE, 2023.\nPradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. A dataset of information-seeking questions and answers anchored in research papers. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4599–4610, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/\nv1/2021.naacl-main.365. URL https://aclanthology.org/2021.naacl-main.365.\nAndrew Davis and Itamar Arel. Low-rank approximations for conditional feedforward computation in deep neural networks, 2014.\n78\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\n\nJeff Dean.\n\nIntroducing Pathways:\n\nA next-generation AI archi-\n\ntecture,\n\n2021.\n\nURL\n\nhttps://blog.google/technology/ai/\n\nintroducing-pathways-next-generation-ai-architecture/.\n\nAparna Dhinakaran, 2024.\n1744771295940669689.\n\nURL https://twitter.com/aparnadhinak/status/\n\nYan Ding, Xiaohan Zhang, Chris Paxton, and Shiqi Zhang. Task and motion planning with large language models for object rearrangement. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2086–2092. IEEE, 2023.\n\nYangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin, Xinyun Chen, Basel Alomair, David Wagner, Baishakhi Ray, and Yizheng Chen. Vulnerability detection with code language models: How far are we? arXiv preprint:2403.18624, 2024.\n\nNan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen MeierHellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V Le, Yonghui Wu, Zhifeng Chen, and Claire Cui. GLaM: Efficient Scaling of Language Models with Mixture-of-Experts. ICML, 2022. URL\nhttps://arxiv.org/abs/2112.06905.\n\nDheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational\nLinguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378,\n2019. URL https://aclanthology.org/N19-1246.\n\nJohn Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12(61):2121–2159, 2011. URL\nhttp://jmlr.org/papers/v12/duchi11a.html.\n\nTyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. Gpts are gpts: An early look at the labor market impact potential of large language models. arXiv preprint arXiv:2303.10130, 2023.\n\nMeng Fang, Xiangpeng Wan, Fei Lu, Fei Xing, and Kai Zou. Mathodyssey: Benchmarking mathematical problem-solving skills in large language models using odyssey math data. arXiv preprint arXiv:2406.18321, 2024.\n\nWilliam Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter\nmodels with simple and efficient sparsity. arXiv preprint arXiv:2101.03961, 2021. URL https: //arxiv.org/abs/2101.03961.\n\nEdward W Felten, Manav Raj, and Robert Seamans. A method to link advances in artificial intelligence to occupational abilities. In AEA Papers and Proceedings, volume 108, pages 54–57. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, 2018.\n\nChrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. Promptbreeder: Self-referential self-improvement via prompt evolution, 2023.\n\nXingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large language models can see but not perceive. arXiv preprint arXiv:2404.12390, 2024.\n\n79\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\n\nXavier Garcia, Yamini Bansal, Colin Cherry, George Foster, Maxim Krikun, Fangxiaoyu Feng, Melvin Johnson, and Orhan Firat. The unreasonable effectiveness of few-shot learning for machine translation, 2023.\n\nSamuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URL\nhttps://aclanthology.org/2020.findings-emnlp.301.\n\nGemini-Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly\ncapable multimodal models. arXiv preprint arXiv:2312.11805, 2023. URL https://storage. googleapis.com/deepmind-media/gemini/gemini_1_report.pdf.\n\nGemma-Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024.\n\nGoogle. Google’s AI Principles.\nprinciples/.\n\n2023.\n\nURL https://ai.google/responsibility/\n\nYash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017.\n\nAlbert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv\npreprint arXiv:2312.00752, 2023. URL https://arxiv.org/abs/2312.00752.\n\nLin Guan, Karthik Valmeekam, Sarath Sreedharan, and Subbarao Kambhampati. Leveraging pretrained large language models to construct and utilize world models for model-based task planning. Advances in Neural Information Processing Systems, 36, 2024.\n\nMandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. Longt5: Efficient text-to-text transformer for long sequences. arXiv preprint arXiv:2112.07916,\n2021. URL https://arxiv.org/abs/2112.07916.\n\nKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. Retrieval augmented language model pre-training. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning\nResearch, pages 3929–3938. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.press/ v119/guu20a.html.\n\nShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. Reasoning with language model is planning with world model. arXiv preprint arXiv:2305.14992, 2023.\n\nJonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner,\nand Marc van Zee. Flax: A neural network library and ecosystem for JAX, 2023. URL http: //github.com/google/flax.\n\nDan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021a.\n\n80\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\n\nDan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. arXiv\npreprint arXiv:2103.03874, 2021b. URL https://arxiv.org/abs/2103.03874.\n\nTom Heskes. On “Natural” Learning and Pruning in Multilayered Perceptrons. Neural Computation,\n12(4):881–901, 04 2000. ISSN 0899-7667. doi: 10.1162/089976600300015637. URL https: //doi.org/10.1162/089976600300015637.\n\nGeoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network, 2015.\n\nJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022. URL\nhttps://arxiv.org/abs/2203.15556.\n\nWenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. Inner monologue: Embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608, 2022.\n\nDaphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A Choquette-Choo, and Nicholas Carlini. Preventing verbatim memorization in language\nmodels gives a false sense of privacy. arXiv preprint arXiv:2210.17546, 2022. URL https:// arxiv.org/abs/210.17546.\n\nGautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. Few-shot learning with retrieval\naugmented language models. arXiv preprint arXiv:2208.03299, 2022. URL https://arxiv.org/ abs/2208.03299.\n\nRobert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. Adaptive mixtures of local experts. Neural computation, 3(1):79–87, 1991.\n\nFrederick Jelinek. Statistical methods for speech recognition. MIT press, 1998.\n\nAlbert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.\n\nZhengbao Jiang, Luyu Gao, Zhiruo Wang, Jun Araki, Haibo Ding, Jamie Callan, and Graham Neubig. Retrieval as attention: End-to-end learning of retrieval and reading within a single transformer. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2336–2349, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.\nemnlp-main.149. URL https://aclanthology.org/2022.emnlp-main.149.\n\nAlex Jones, Isaac Caswell, Ishank Saxena, and Orhan Firat. Bilex rx: Lexical data augmentation for massively multilingual machine translation. arXiv preprint arXiv:2303.15265, 2023.\n\nRafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. Exploring the limits\nof language modeling, 2016. URL https://arxiv.org/abs/1602.02410.\n\nGregory Kamradt, 2023.\n\nURL https://github.com/gkamradt/LLMTest_\n\nNeedleInAHaystack/blob/main/README.md.\n\n81\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nJared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv\npreprint arXiv:2001.08361, 2020. URL https://arxiv.org/abs/2001.08361.\nVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.550. URL\nhttps://aclanthology.org/2020.emnlp-main.550.\nAniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In ECCV, 2016.\nYoung Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla. Scalable and efficient moe training for multitask multilingual models. arXiv preprint arXiv:2109.10465, 2021.\nMegan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R Lin, Hjalmar Wijk, Joel Burget, et al. Evaluating language-model agents on realistic autonomous tasks. arXiv preprint arXiv:2312.11671, 2023.\nAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023.\nR. Kneser and H. Ney. Improved backing-off for m-gram language modeling. In 1995 International Conference on Acoustics, Speech, and Signal Processing, volume 1, pages 181–184 vol.1, 1995. doi: 10.1109/ICASSP.1995.479394.\nTom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Makoto Nagata, Toshiaki Nakazawa, Martin Popel, Maja Popović, and Mariya Shmatova. Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet. In Philipp Koehn, Barry Haddow, Tom Kocmi, and Christof Monz, editors, Proceedings of the Eighth Conference on Machine Translation, pages 1–42, Singapore, December 2023. Association for Computational\nLinguistics. doi: 10.18653/v1/2023.wmt-1.1. URL https://aclanthology.org/2023.wmt-1. 1.\nGreg Kohs. Alphago. Motion Picture, 2017. Produced by DeepMind Technologies and distributed by Netflix.\nSneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. Madlad-400: A multilingual and document-level large audited dataset, 2023.\nJ. Landeghem, R. Powalski, R. Tito, D. Jurkiewicz, M. Blaschko, L. Borchmann, M. Coustaty, S. Moens, M. Pietruszka, B. Ackaert, T. Stanislawek, P. Joziak, and E. Valveny. Document understanding dataset and evaluation (dude). In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 19471–19483, 2023.\n82\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. GShard: Scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations, 2020. URL\nhttps://openreview.net/forum?id=qrwe7XHTmYb.\nAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. Solving quantitative reasoning problems with language models. arXiv preprint arXiv:2206.14858, 2022. URL\nhttps://arxiv.org/abs/2206.14858.\nFangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, and Julian Eisenschlos. MatCha: Enhancing visual language pretraining with math reasoning and chart derendering. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12756–12770, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.714. URL\nhttps://aclanthology.org/2023.acl-long.714.\nHao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. World model on million-length video and language with ringattention, 2024.\nPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, KaiWei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023.\nMAA. American invitational mathematics examination - aime. In American Invitational Mathemat-\nics Examination - AIME 2024, February 2024. URL https://maa.org/math-competitions/ american-invitational-mathematics-examination-aime.\nMacknight, Aung, and Gomes. Personal Communication.\nArjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, et al. Openeqa: Embodied question answering in the era of foundation models. In 2nd Workshop on Mobile Manipulation and Embodied Intelligence at ICRA 2024, 2024.\nChaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das, Mirella Lapata, and Chris Alberti. Dolomites: Domain-specific long-form methodical tasks, 2024.\nKarttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. EgoSchema: A diagnostic benchmark for very long-form video language understanding. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023.\nPedro Henrique Martins, Zita Marinho, and Andre Martins. ∞-former: Infinite memory transformer. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5468–5485, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/\nv1/2022.acl-long.375. URL https://aclanthology.org/2022.acl-long.375.\nAhmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of ACL, 2022.\n83\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nMATH-AI-2023-Panel. Panel discussion. In MATH-AI Workshop at NeurIPS 2023, New Orleans,\nLouisiana, USA, December 2023. URL https://mathai2023.github.io/.\nMinesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2200–2209, 2021.\nMinesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706, 2022.\nMantas Mazeika, Andy Zou, Norman Mu, Long Phan, Zifan Wang, Chunru Yu, Adam Khoja, Fengqing Jiang, Aidan O’Gara, Ellie Sakhaee, Zhen Xiang, Arezoo Rajabi, Dan Hendrycks, Radha Poovendran, Bo Li, and David Forsyth. Tdc 2023 (llm edition): The trojan detection challenge. In NeurIPS Competition Track, 2023.\nMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249, 2024.\nJ. McMurry. Organic Chemistry. International edition. Brooks/Cole Cengage Learning, 2012. ISBN\n9780840054531. URL https://books.google.at/books?id=oVv4Az7VJRYC.\nTomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernock`y, and Sanjeev Khudanpur. Recurrent neural network based language model. In INTERSPEECH, 2010.\nSewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11048–11064, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.\nemnlp-main.759. URL https://aclanthology.org/2022.emnlp-main.759.\nMargaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, page 220–229, New York, NY, USA, 2019a. Association for Computing Machinery. ISBN 9781450361255. doi:\n10.1145/3287560.3287596. URL https://doi.org/10.1145/3287560.3287596.\nMargaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the conference on Fairness, Accountability, and Transparency, pages 220–229, 2019b.\nJesse Mu, Xiang Lisa Li, and Noah Goodman. Learning to compress prompts with gist tokens. arXiv preprint arXiv:2304.08467, 2023.\nMilad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. Scalable extraction of training data from (production) language models, 2023.\nOpenAI. GPT-4 Technical Report. 2023a.\nOpenAI. GPT-4V(ision) System Card, 2023b.\n84\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nOpenAI. Whisper, 2023. URL https://github.com/openai/whisper.\nAntonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De. Resurrecting recurrent neural networks for long sequences. arXiv preprint arXiv:2303.06349, 2023.\nAlicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman. BBQ: A hand-built bias benchmark for question answering.\nCoRR, abs/2110.08193, 2021. URL https://arxiv.org/abs/2110.08193.\nMary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe, and Toby Shevlane. Evaluating frontier models for dangerous capabilities. arXiv preprint:2403.13793, 2024.\nReiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. Efficiently scaling transformer inference. Proceedings of Machine Learning and Systems, 5, 2023.\nMaja Popović. chrF: character n-gram F-score for automatic MT evaluation. In Ondřej Bojar, Rajan Chatterjee, Christian Federmann, Barry Haddow, Chris Hokamp, Matthias Huck, Varvara Logacheva, and Pavel Pecina, editors, Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi:\n10.18653/v1/W15-3049. URL https://aclanthology.org/W15-3049.\nVineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert. Mls: A large-scale multilingual dataset for speech research. arXiv preprint arXiv:2012.03411, 2020.\nOfir Press, Noah A Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021.\nJack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, JeanBaptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training Gopher. CoRR, abs/2112.11446, 2021.\nColin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-\nto-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http: //jmlr.org/papers/v21/20-074.html.\n85\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nDavid Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022, 2023.\nCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby. Scaling vision with sparse mixture of experts, 2021.\nWilliam A Gaviria Rojas, Sudnya Diamos, Keertan Ranjan Kini, David Kanter, Vijay Janapa Reddi, and Cody Coleman. The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022.\nStephen Roller, Sainbayar Sukhbaatar, Jason Weston, et al. Hash layers for large sparse models. Advances in Neural Information Processing Systems, 34:17555–17566, 2021.\nStuart J Russell and Peter Norvig. Artificial intelligence: a modern approach. Pearson, 2016.\nKhaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, et al. Capabilities of Gemini models in medicine. arXiv preprint arXiv:2404.18416, 2024.\nKeshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. Colbertv2: Effective and efficient retrieval via lightweight late interaction. arXiv preprint arXiv:2112.01488, 2021.\nThibault Sellam, Dipanjan Das, and Ankur Parikh. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.\nacl-main.704. URL https://aclanthology.org/2020.acl-main.704.\nClaude Elwood Shannon. A mathematical theory of communication. The Bell System Technical Jour-\nnal, 27:379–423, 1948. URL http://plan9.bell-labs.com/cm/ms/what/shannonday/ shannon1948.pdf.\nNoam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.\nIn ICLR (Poster). OpenReview.net, 2017. URL http://dblp.uni-trier.de/db/conf/iclr/ iclr2017.html#ShazeerMMDLHD17.\nToby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, and Allan Dafoe. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324, 2023.\nFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language Models\nare Multilingual Chain-of-Thought Reasoners. In Proceedings of ICLR 2023, 2023a. URL http: //arxiv.org/abs/2210.03057.\nWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Victoria Lin, Noah A Smith, Luke Zettlemoyer, Scott Yih, and Mike Lewis. In-context pretraining: Language modeling beyond document boundaries. arXiv preprint arXiv:2310.10638, 2023b.\n86\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nAmanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards VQA models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019.\nIshika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 11523–11530. IEEE, 2023.\nAarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities\nof language models. arXiv preprint arXiv:2206.04615, 2022. URL https://arxiv.org/abs/ 2206.04615.\nSaurabh Srivastava, Annarose M B, Anto P V au2, Shashank Menon, Ajay Sukumar, Adwaith Samod T, Alan Philipose, Stevin Prince, and Sooraj Thomas. Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap, 2024.\nKonrad Staniszewski, Szymon Tworkowski, Sebastian Jaszczur, Henryk Michalewski, Łukasz Kuciński, and Piotr Miłoś. Structured packing in llm training improves long context utilization. arXiv preprint arXiv:2312.17296, 2023.\nMirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.\nGarrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky, and Luke Melas-Kyriazi. A benchmark for learning to translate a new language from one grammar book. In Arxiv, 2023.\nNLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. No language left behind: Scaling human-centered machine translation. 2022.\nRomal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. LaMDA: Language models for dialog applications.\narXiv preprint arXiv:2201.08239, 2022. URL https://arxiv.org/abs/2201.08239.\nKocmi Tom, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, et al. Findings of the 2023 conference on machine translation (wmt23): Llms are here but not quite there yet. In WMT23-Eighth Conference on Machine Translation, pages 198–216, 2023.\nShubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman. Openmathinstruct-1: A 1.8 million math instruction tuning dataset, 2024.\nHugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.\n87\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nHugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.\nKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati. On the planning abilities of large language models-a critical investigation. Advances in Neural Information Processing Systems, 36, 2024.\nAshish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. CoRR, abs/1706.03762, 2017. URL\nhttp://arxiv.org/abs/1706.03762.\nRamakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4566–4575, 2015.\nEline Visser. Kalamang dictionary. Dictionaria, (13):1–2737, 2020a. URL https://dictionaria. clld.org/contributions/kalamang.\nEline Visser. A grammar of kalamang: The papuan language of the karas islands. 2020b.\nEline Visser. The Kalamang collection: an archive of linguistic and cultural material from Karas.\nLund: Humlab, Lund University Humanities Lab, 2020c. URL http://hdl.handle.net/10050/ 00-0000-0000-0003-C3E8-1.\nChanghan Wang, Anne Wu, and Juan Pino. Covost 2 and massively multilingual speech-to-text translation. arXiv preprint arXiv:2007.10310, 2020.\nChanghan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. arXiv preprint arXiv:2101.00390, 2021.\nXin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. VATEX: A large-scale, high-quality multilingual dataset for video-and-language research. In ICCV, 2019a.\nXinda Wang, Kun Sun, Archer Batcheller, and Sushil Jajodia. Detecting \"0-day\" vulnerability: An empirical study of secret security patch in OSS. In IEEE/IFIP International Conference on Depend-\nable Systems and Networks, pages 485–492. IEEE, 2019b. URL https://csis.gmu.edu/ksun/ publications/secretpatch-dsn19.pdf.\nLaura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359, 2021.\n88\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\n\nLaura Weidinger, Joslyn Barnhart, Jenny Brennan, Christina Butterfield, Susie Young, Will Hawkins, Lisa Anne Hendricks, Ramona Comanescu, Oscar Chang, Mikel Rodriguez, et al. Holistic safety and responsibility evaluations of advanced ai models. arXiv preprint arXiv:2404.14068, 2024.\n\nWikipedia contributors. Skeletal formula — Wikipedia, the free encyclopedia, 2024. URL https:// en.wikipedia.org/w/index.php?title=Skeletal_formula&oldid=1221781655. [On-\nline; accessed 7-May-2024].\n\nWorld Economic Forum. Jobs of Tomorrow: Large Language Models and Jobs. 2023. URL https: //www3.weforum.org/docs/WEF_Jobs_of_Tomorrow_Generative_AI_2023.pdf.\n\nPenghao Wu and Saining Xie. V*: Guided visual search as a core mechanism in multimodal llms. arXiv preprint arXiv:2312.14135, 2023.\n\nQingyang Wu, Zhenzhong Lan, Kun Qian, Jing Gu, Alborz Geramifard, and Zhou Yu. Memformer: A memory-augmented transformer for sequence modeling. In Yulan He, Heng Ji, Sujian Li, Yang Liu, and Chua-Hui Chang, editors, Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, pages 308–318, Online only, November 2022a. Association for Computational Linguistics.\nURL https://aclanthology.org/2022.findings-aacl.29.\n\nYuhuai Wu, Markus N Rabe, DeLesley Hutchins, and Christian Szegedy. Memorizing transformers. arXiv preprint arXiv:2203.08913, 2022b.\n\nwunderwuzzi. Hacking Google Bard - From Prompt Injection to Data Exfiltration. https://embracethered.com/blog/posts/2023/google-bard-data-exfiltration/, 2023.\n\nx.ai. Grok-1.5 vision preview. URL https://x.ai/blog/grok-1.5v.\n\nWenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al. Effective long-context scaling of foundation models. arXiv preprint arXiv:2309.16039, 2023.\n\nXLA. XLA: Optimizing compiler for TensorFlow. https://www.tensorflow.org/xla, 2019.\n[Online; accessed December-2023].\n\nYuanzhong Xu, HyoukJoong Lee, Dehao Chen, Blake Hechtman, Yanping Huang, Rahul Joshi, Maxim Krikun, Dmitry Lepikhin, Andy Ly, Marcello Maggioni, et al. Gspmd: general and scalable parallelization for ml computation graphs. arXiv preprint arXiv:2105.04663, 2021.\n\nFanjia Yan, Huanzhi Mao, Charlie Cheng-Jie Ji, Tianjun Zhang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. Berkeley function calling leaderboard. 2024.\n\nChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. Large language models as optimizers, 2023a.\n\nJohn Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. InterCode: Standardizing\nand benchmarking interactive coding with execution feedback. ArXiv, June 2023b. URL https: //arxiv.org/abs/2306.14898.\n\nKenneth Yeung.\n\nNew google gemini vulnerability enabling pro-\n\nfound misuse.\n\n2024.\n\nURL https://hiddenlayer.com/\n\nresearch/new-google-gemini-content-manipulation-vulns-found/\n\n#Indirect-Injections-are-back!\n\n89\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nEmine Yilmaz, Manisha Verma, Rishabh Mehrotra, Evangelos Kanoulas, Ben Carterette, and Nick Craswell. Overview of the TREC 2015 tasks track. In Proceedings of the Twenty-Fourth Text RE-\ntrieval Conference (TREC 2015), 2015. URL https://trec.nist.gov/pubs/trec24/papers/ Overview-T.pdf.\nZhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. ActivityNet-QA: A dataset for understanding complex web videos via question answering. In AAAI, 2019.\nXiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2023.\nManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. Big bird: Transformers for longer sequences. Advances in Neural Information Processing Systems, 33:17283–17297, 2020.\nRowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.\nBiao Zhang, Barry Haddow, and Alexandra Birch. Prompting large language model for machine translation: a case study. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023a.\nMichael Zhang and Christopher Ré. Contrastive adapters for foundation model group robustness. Advances in Neural Information Processing Systems, 35:21682–21697, 2022.\nYu Zhang, Wei Han, James Qin, Yongqiang Wang, Ankur Bapna, Zhehuai Chen, Nanxin Chen, Bo Li, Vera Axelrod, Gary Wang, Zhong Meng, Ke Hu, Andrew Rosenberg, Rohit Prabhavalkar, Daniel S. Park, Parisa Haghani, Jason Riesa, Ginger Perng, Hagen Soltau, Trevor Strohman, Bhuvana Ramabhadran, Tara Sainath, Pedro Moreno, Chung-Cheng Chiu, Johan Schalkwyk, Françoise Beaufays, and Yonghui Wu. Google usm: Scaling automatic speech recognition beyond 100 languages. arXiv preprint arXiv:2303.01037, 2023b.\nDora Zhao, Angelina Wang, and Olga Russakovsky. Understanding and evaluating racial biases in image captioning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14830–14840, 2021.\nZexuan Zhong, Tao Lei, and Danqi Chen. Training language models with memory augmentation. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5657–5673, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.\nemnlp-main.382. URL https://aclanthology.org/2022.emnlp-main.382.\nLuowei Zhou, Chenliang Xu, and Jason J Corso. Towards automatic learning of procedures from web instructional videos. In AAAI Conference on Artificial Intelligence, pages 7590–7598, 2018.\nYaqin Zhou, Jing Kai Siow, Chenyu Wang, Shangqing Liu, and Yang Liu. SPI: Automated identification of security patches via commits. ACM Transactions on Software Engineering and Methodology, 31 (1):1–27, 2021.\nFengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. Towards complex document understanding by discrete reasoning. In Proceedings of the 30th ACM International Conference on Multimedia, page 4857–4866, 2022.\n90\n\n\fGemini 1.5: Unlocking multimodal understanding across millions of tokens of context\nBarret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. Designing effective sparse expert models. arXiv preprint arXiv:2202.08906, 2022.\nURL https://arxiv.org/abs/2202.08906.\nAndy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    },
    {
      "Slug": "deepseek-r1",
      "Paper": "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning",
      "AtlasYear": 2025,
      "Status": "indexed",
      "Method": "native-pdf-text",
      "SourceUrl": "https://arxiv.org/pdf/2501.12948v1",
      "PdfSha256": "52D8CA3AC93E88CEF9944E1FD03B0E04AEC5954495A8250FB2FADF8FA20A4DAD",
      "Sections": [
        {
          "Section": "Reference section 1",
          "StartLine": 753,
          "EndLine": 809,
          "PdfPages": [
            17,
            18,
            19
          ],
          "Text": "AI@Meta. Llama 3.1 model card, 2024. URL https://github.com/meta-llama/llama-m odels/blob/main/models/llama3_1/MODEL_CARD.md.\nAnthropic. Claude 3.5 sonnet, 2024. URL https://www.anthropic.com/news/claude-3 -5-sonnet.\nM. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021.\nURL https://arxiv.org/abs/2107.03374.\nA. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.\nY. Dubois, B. Galambosi, P. Liang, and T. B. Hashimoto. Length-controlled alpacaeval: A simple way to debias automatic evaluators. arXiv preprint arXiv:2404.04475, 2024.\nX. Feng, Z. Wan, M. Wen, S. M. McAleer, Y. Wen, W. Zhang, and J. Wang. Alphazero-like\ntree-search can guide large language model decoding and training, 2024. URL https: //arxiv.org/abs/2309.17179.\nL. Gao, J. Schulman, and J. Hilton. Scaling laws for reward model overoptimization, 2022. URL\nhttps://arxiv.org/abs/2210.10760.\nA. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y. Zhao, X. Du, M. R. G. Madani, C. Barale, R. McHardy, J. Harris, J. Kaddour, E. van Krieken, and\nP. Minervini. Are we done with mmlu? CoRR, abs/2406.04127, 2024. URL https://doi.or g/10.48550/arXiv.2406.04127.\nGoogle. Our next-generation model: Gemini 1.5, 2024. URL https://blog.google/techno logy/ai/google-gemini-next-generation-model-february-2024.\nY. He, S. Li, J. Liu, Y. Tan, W. Wang, H. Huang, X. Bu, H. Guo, C. Hu, B. Zheng, et al. Chinese simpleqa: A chinese factuality evaluation for large language models. arXiv preprint arXiv:2411.07140, 2024.\nD. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.\nY. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, et al. C-Eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322, 2023.\nN. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code.\nCoRR, abs/2403.07974, 2024. URL https://doi.org/10.48550/arXiv.2403.07974.\n17\n\n\fS. Krishna, K. Krishna, A. Mohananey, S. Schwarcz, A. Stambler, S. Upadhyay, and M. Faruqui. Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation. CoRR,\nabs/2409.12941, 2024. doi: 10.48550/ARXIV.2409.12941. URL https://doi.org/10.485 50/arXiv.2409.12941.\nA. Kumar, V. Zhuang, R. Agarwal, Y. Su, J. D. Co-Reyes, A. Singh, K. Baumli, S. Iqbal, C. Bishop, R. Roelofs, et al. Training language models to self-correct via reinforcement learning. arXiv preprint arXiv:2409.12917, 2024.\nH. Li, Y. Zhang, F. Koto, Y. Yang, H. Zhao, Y. Gong, N. Duan, and T. Baldwin. CMMLU: Measuring massive multitask language understanding in Chinese. arXiv preprint arXiv:2306.09212, 2023.\nT. Li, W.-L. Chiang, E. Frick, L. Dunlap, T. Wu, B. Zhu, J. E. Gonzalez, and I. Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. arXiv preprint arXiv:2406.11939, 2024.\nH. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.\nB. Y. Lin. ZeroEval: A Unified Framework for Evaluating Language Models, July 2024. URL\nhttps://github.com/WildEval/ZeroEval.\nMAA. American invitational mathematics examination - aime. In American Invitational\nMathematics Examination - AIME 2024, February 2024. URL https://maa.org/math -competitions/american-invitational-mathematics-examination-aime.\nOpenAI. Hello GPT-4o, 2024a. URL https://openai.com/index/hello-gpt-4o/.\nOpenAI. Learning to reason with llms, 2024b. URL https://openai.com/index/learnin g-to-reason-with-llms/.\nOpenAI. Introducing SimpleQA, 2024c. URL https://openai.com/index/introducing -simpleqa/.\nOpenAI. Introducing SWE-bench verified we’re releasing a human-validated subset of swe-\nbench that more, 2024d. URL https://openai.com/index/introducing-swe-bench -verified/.\nQwen. Qwq: Reflect deeply on the boundaries of the unknown, 2024a. URL https://qwenlm .github.io/blog/qwq-32b-preview/.\nQwen. Qwen2.5: A party of foundation models, 2024b. URL https://qwenlm.github.io/b log/qwen2.5.\nD. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022, 2023.\nZ. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. Li, Y. Wu, and D. Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.\nD. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. P. Lillicrap, K. Simonyan, and D. Hassabis. Mastering chess and shogi by self-play with a general reinforcement learning algorithm. CoRR, abs/1712.01815,\n2017a. URL http://arxiv.org/abs/1712.01815.\n18\n\n\fD. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. P. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis. Mastering the game of go without human knowledge. Nat., 550(7676):354–359,\n2017b. doi: 10.1038/NATURE24270. URL https://doi.org/10.1038/nature24270.\nC. Snell, J. Lee, K. Xu, and A. Kumar. Scaling llm test-time compute optimally can be more\neffective than scaling model parameters, 2024. URL https://arxiv.org/abs/2408.033 14.\nT. Trinh, Y. Wu, Q. Le, H. He, and T. Luong. Solving olympiad geometry without human demonstrations. Nature, 2024. doi: 10.1038/s41586-023-06747-5.\nJ. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022.\nP. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui. Math-shepherd: A labelfree step-by-step verifier for llms in mathematical reasoning. arXiv preprint arXiv:2312.08935, 2023.\nX. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.\nY. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. CoRR, abs/2406.01574, 2024.\nURL https://doi.org/10.48550/arXiv.2406.01574.\nC. S. Xia, Y. Deng, S. Dunn, and L. Zhang. Agentless: Demystifying llm-based software engineering agents. arXiv preprint, 2024.\nH. Xin, Z. Z. Ren, J. Song, Z. Shao, W. Zhao, H. Wang, B. Liu, L. Zhang, X. Lu, Q. Du, W. Gao, Q. Zhu, D. Yang, Z. Gou, Z. F. Wu, F. Luo, and C. Ruan. Deepseek-prover-v1.5: Harnessing proof assistant feedback for reinforcement learning and monte-carlo tree search, 2024. URL\nhttps://arxiv.org/abs/2408.08152.\nJ. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023."
        }
      ],
      "PublisherReferences": [],
      "Notes": [
        "Raw extraction retains OCR errors, ligatures, page headers and column-order artifacts. Entries have not all been normalized or individually verified."
      ]
    }
  ]
}
